EDBT 2026 Demo / reviewers in the wild / expert
Mengke Li 0001
dblp:67/6597-1
· DBLP profile ↗
37ranked-venue papers
8as first author
35since 2021 · last 2026
0000-0002-9433-9683ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 4 first-author · 20 since 2021Artificial intelligence and machine learning · 19 · 6 first-author · 19 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Break the Tie: Learning Cluster-Customized Category Relationships for Categorical Data ClusteringabstractCategorical attributes with qualitative values are ubiquitous in cluster analysis of real datasets. Unlike the Euclidean distance of numerical attributes, the categorical attributes lack well-defined relationships of their possible values (also called categories interchangeably), which hampers the exploration of compact categorical data clusters. Although most attempts are made for developing appropriate distance metrics, they typically assume a fixed topological relationship between categories when learning distance metrics, which limits their adaptability to varying cluster structures and often leads to suboptimal clustering performance. This paper, therefore, breaks the intrinsic relationship tie of attribute categories and learns customized distance metrics suitable for flexibly and accurately revealing various cluster distributions. As a result, the fitting ability of the clustering algorithm is significantly enhanced, benefiting from the learnable category relationships. Moreover, the learned category relationships are proved to be Euclidean distance metric-compatible, enabling a seamless extension to mixed datasets that include both numerical and categorical attributes. Comparative experiments on 12 real benchmark datasets with significance tests show the superior clustering accuracy of the proposed method with an average ranking of 1.25, which is significantly higher than the 5.21 ranking of the best-performing methods. Code and extended version with detailed proofs are provided online. Mingjie Zhao 0003, Zhanpei Huang, Yang Lu 0009, Mengke Li 0001, Yiqun Zhang 0006, Weifeng Su, Yiu-Ming Cheung |
AAAI | 4 |
| 2026 | HyReaL: Clustering Attributed Graph via Hyper-complex Space Representation Learning
Yang Lu 0009, Mengke Li 0001, Cuie Yang, Yiqun Zhang 0006, Yiu-Ming Cheung |
DASFAA (2) | 3 |
| 2025 | Asynchronous Federated Clustering with Unknown Number of ClustersabstractFederated Clustering (FC) is crucial to mining knowledge from unlabeled non-Independent Identically Distributed (non-IID) data provided by multiple clients while preserving their privacy. Most existing attempts learn cluster distributions at local clients, then securely pass the desensitized information to the server for aggregation. However, some tricky but common FC problems are still relatively unexplored, including the heterogeneity in terms of clients' communication capacity and the unknown number of proper clusters. To further bridge the gap between FC and real application scenarios, this paper first shows that the clients' communication asynchrony and unknown proper cluster numbers are complex coupling problems, and then proposes an Asynchronous Federated Cluster Learning (AFCL) method accordingly. It spreads the excessive number of seed points to clients as a learning medium and coordinates them across clients to form a consensus. To alleviate the distribution imbalance cumulated due to the unforeseen asynchronous uploading from the heterogeneous clients, we also design a balancing mechanism for seeds updating. As a result, the seeds gradually adapt to each other to reveal a proper number of clusters. Extensive experiments demonstrate the efficacy of AFCL. Yiqun Zhang 0006, Yang Lu 0009, Mengke Li 0001, Yiu-Ming Cheung |
AAAI | 4 |
| 2025 | Weighted Density for The Win: Accurate Subspace Density Clusteringabstractk-clustering typically struggles with the detection of irregular-distributed clusters due to the natural bias, while density clustering usually cannot well-adapt to different datasets and clustering tasks as it is not an oriented optimization process. This paper, therefore, proposes to perform density clustering in dynamically learned subspaces. To exploit the irregular-distributed clusters obtained by density clustering for the subspace determination, we design a new strategy to appropriately evaluate the importance of attributes. It turns out that the proposed Weighted Density-based Subspace Clustering (WDSC) algorithm inherits the unbiased merits of density clustering, and also upgrades the unlearning density clustering to be learnable under the subspace learning paradigm of k-clustering. A comprehensive evaluation including significance tests, ablation studies, qualitative comparisons, etc., shows the superiority of WDSC. Maixuan Peng, Yuyang Wu, Yang Lu 0009, Mengke Li 0001, Yiqun Zhang 0006, Yiu-Ming Cheung |
ICASSP | 4 |
| 2025 | PRO-VPT: Distribution-Adaptive Visual Prompt Tuning via Prompt RelocationabstractVisual prompt tuning (VPT), i.e., fine-tuning some lightweight prompt tokens, provides an efficient and effective approach for adapting pre-trained models to various downstream tasks. However, most prior art indiscriminately uses a fixed prompt distribution across different tasks, neglecting the importance of each block varying depending on the task. In this paper, we introduce adaptive distribution optimization (ADO) by tackling two key questions: (1) How to appropriately and formally define ADO, and (2) How to design an adaptive distribution strategy guided by this definition? Through empirical analysis, we first confirm that properly adjusting the distribution significantly improves VPT performance, and further uncover a key insight that a nested relationship exists between ADO and VPT. Based on these findings, we propose a new VPT framework, termed PRO-VPT (iterative Prompt RelOcation-based VPT), which adaptively adjusts the distribution built upon a nested optimization formulation. Specifically, we develop a prompt relocation strategy derived from this formulation, comprising two steps: pruning idle prompts from prompt-saturated blocks, followed by allocating these prompts to the most prompt-needed blocks. By iteratively performing prompt relocation and VPT, our proposal can adaptively learn the optimal prompt distribution in a nested optimization-based manner, thereby unlocking the full potential of VPT. Extensive experiments demonstrate that our proposal significantly outperforms advanced VPT methods, e.g., PRO-VPT surpasses VPT by 1.6 pp and 2.0 pp average accuracy, leading prompt-based methods to state-of-the-art performance on VTAB-1k and FGVC benchmarks. The code is available at https://github.com/ckshang/PRO-VPT. Chikai Shang, Mengke Li 0001, Yiqun Zhang 0006, Zhen Chen 0018, Jinlin Wu, Fangqing Gu, Yang Lu 0009, Yiu-Ming Cheung |
ICCV | 2 |
| 2025 | Gamma Distribution PCA-Enhanced Feature Learning for Angle-Robust SAR Target RecognitionabstractScattering characteristics of synthetic aperture radar (SAR) targets are typically related to observed azimuth and depression angles. However, in practice, it is difficult to obtain adequate training samples at all observation angles, which probably leads to poor robustness of deep networks. In this paper, we first propose a Gamma-Distribution Principal Component Analysis ($\Gamma$PCA) model that fully accounts for the statistical characteristics of SAR data. The $\Gamma$PCA derives consistent convolution kernels to effectively capture the angle-invariant features of the same target at various attitude angles, thus alleviating deep models’ sensitivity to angle changes in SAR target recognition task. We validate $\Gamma$PCA model based on two commonly used backbones, ResNet and ViT, and conduct multiple robustness experiments on the MSTAR benchmark dataset. The experimental results demonstrate that $\Gamma$PCA effectively enables the model to withstand substantial distributional discrepancy caused by angle changes. Additionally, $\Gamma$PCA convolution kernel is designed to require no parameter updates, introducing no extra computational burden to the network. The source code is available at https://github.com/ChGrey/GammaPCA. Mengke Li 0001 |
ICML | 3 |
| 2025 | FATE: A Prompt-Tuning-Based Semi-Supervised Learning Framework for Extremely Limited Labeled DataabstractSemi-supervised learning (SSL) has achieved significant progress by leveraging both labeled data and unlabeled data. Existing SSL methods overlook a common real-world scenario when labeled data is extremely scarce, potentially as limited as a single labeled sample in the dataset. General SSL approaches struggle to train effectively from scratch under such constraints, while methods utilizing pre-trained models often fail to find an optimal balance between leveraging limited labeled data and abundant unlabeled data. To address this challenge, we propose Firstly Adapt, Then catEgorize (FATE), a novel SSL framework tailored for scenarios with extremely limited labeled data. At its core, the two-stage prompt tuning paradigm FATE exploits unlabeled data to compensate for scarce supervision signals, then transfers to downstream tasks. Concretely, FATE first adapts a pre-trained model to the feature distribution of downstream data using volumes of unlabeled samples in an unsupervised manner. It then applies an SSL method specifically designed for pre-trained models to complete the final classification task. FATE is designed to be compatible with both vision and vision-language pre-trained models. Extensive experiments demonstrate that FATE effectively mitigates challenges arising from the scarcity of labeled samples in SSL, achieving an average performance improvement of 33.74% across seven benchmarks compared to state-of-the-art SSL methods. Code is available at https://github.com/ganchi-huanggua/FATE.git. Hezhao Liu, Yang Lu 0009, Mengke Li 0001, Yiqun Zhang 0006, Shreyank N. Gowda, Chen Gong 0002, Hanzi Wang |
ACM Multimedia | 3 |
| 2025 | Unlocker: Disentangle the Deadlock of Learning between Label-noisy and Long-tailed DataabstractIn real world, the observed label distribution of a dataset often mismatches its true distribution due to noisy labels.
In this situation, noisy labels learning (NLL) methods directly integrated with long-tail learning (LTL) methods tend to fail due to a dilemma: NLL methods normally rely on unbiased model predictions to recover true distribution by selecting and correcting noisy labels; while LTL methods like logit adjustment depends on true distributions to adjust biased predictions, leading to a deadlock of mutual dependency defined in this paper.
To address this, we propose \texttt{Unlocker}, a bilevel optimization framework that integrates NLL methods and LTL methods to iteratively disentangle this deadlock. The inner optimization leverages NLL to train the model, incorporating LTL methods to fairly select and correct noisy labels. The outer optimization adaptively determines an adjustment strength, mitigating model bias from over- or under-adjustment. We also theoretically prove that this bilevel optimization problem is convergent by transferring the outer optimization target to an equivalent problem with a closed-form solution.
Extensive experiments on synthetic and real-world datasets demonstrate the effectiveness of our method in alleviating model bias and handling long-tailed noisy label data. Code is available at \url{https://anonymous.4open.science/r/neurips-2025-anonymous-1015/}. Chen Shu, Ruichi Zhang, Mengke Li 0001, Yonggang Zhang 0003, Yang Lu 0009, Bo Han 0003, Yiu-Ming Cheung, Hanzi Wang |
NeurIPS | 4 |
| 2025 | ESCOR: Emotion-Aware Semantic Constraint and Correlation Refinement for Image Emotion Distribution Learning
Haotian Wu 0009, Mengke Li 0001, Yiu-Ming Cheung, Zhihong Tian 0001 |
PRCV (12) | 3 |
| 2025 | Data Enhancement for Long-tailed Tasks: Diffusion Model with Optimized Quality FilterabstractIn the field of long-tail learning, the uneven distribution of datasets leads to a significant decrease in the model’s accuracy for tail classes. The data augmentation method is an effective way to alleviate the long-tail problem. The diversity of most existing augmentation methods is insufficient, leading to the limited improvement of tail class information. Due to the diffusion model’s ability to generate images with high quality, rich details, and diversity by gradually denoising images, we propose a augmentation method based on the diffusion model called Diffusion-VagMix (DVM). Based on the diffusion model, this method uses an Optimized Quality Filter (OptiFilter) to scalp low-quality generated images and perform further data augmentation. Our approach effectively addresses the issue of insufficient samples in tail classes and uneven quality of the resulting images by augmenting the dataset with high-quality generated images. This strategy leads to improvements in the accuracy of tail classes and enhances the model’s overall performance. Our DVM method can be used on other long-tail learning methods at will, which can make further improvements. The source code is available at https://github.com/Alert-M/Diffusion-VagMix. Jiawei You, Yichuan Zhai, Mengke Li 0001, Yang Lu 0009 |
SMC | 3 |
| 2025 | Categorical Data Clustering via Value Order Estimated Distance Metric LearningabstractClustering is a popular machine learning technique for data mining that can process and analyze datasets to automatically reveal sample distribution patterns. Since the ubiquitous categorical data naturally lack a well-defined metric space such as the Euclidean distance space of numerical data, the distribution of categorical data is usually under-represented, and thus valuable information can be easily twisted in clustering. This paper, therefore, introduces a novel order distance metric learning approach to intuitively represent categorical attribute values by learning their optimal order relationship and quantifying their distance in a line similar to that of the numerical attributes. Since subjectively created qualitative categorical values involve ambiguity and fuzziness, the order distance metric is learned in the context of clustering. Accordingly, a new joint learning paradigm is developed to alternatively perform clustering and order distance metric learning with low time complexity and a guarantee of convergence. Due to the clustering-friendly order learning mechanism and the homogeneous ordinal nature of the order distance and Euclidean distance, the proposed method achieves superior clustering accuracy on categorical and mixed datasets. More importantly, the learned order distance metric greatly reduces the difficulty of understanding and managing the non-intuitive categorical data. Experiments with ablation studies, significance tests, case studies, etc., have validated the efficacy of the proposed method. The source code is available at https://github.com/csmjzhao/OCL_Source_Code. Yiqun Zhang 0006, Mingjie Zhao 0003, Hong Jia, Mengke Li 0001, Yang Lu 0009, Yiu-Ming Cheung |
Proc. ACM Manag. Data | 4 |
| 2025 | MOOD: Leveraging Out-of-Distribution Data to Enhance Imbalanced Semi-Supervised LearningabstractThe imbalanced semi-supervised learning (SSL) has emerged as a critical research area due to the prevalence of class imbalanced and partially labeled data in real-world scenarios. As the requirement for data volume increases, naturally collected datasets inevitably contain out-of-distribution (OOD) samples. However, the performance of existing imbalanced SSL methods experiences a marked deterioration with OOD data. In this article, we propose an imbalanced SSL method called mixup-OOD (MOOD) to address this issue. The core idea is to "turn waste into treasure," exploring the potential of leveraging seemingly detrimental OOD data to expand the feature space, particularly for tail classes. Specifically, we first filter OOD data from unlabeled data, and then fuse it with labeled data to boost feature diversity for the tail classes. To avoid feature overlapping with OOD data, we develop a push-and-pull (PaP) loss to attract in-distribution (ID) instances toward respective class centroids while repelling OOD samples from them. Extensive experiments show that MOOD achieves superior performance compared with other state-of-the-art methods and exhibits robustness across data with different imbalanced ratios and OOD proportions. The source code is available at: https://github.com/xlhuang132/MOODv2. Yang Lu 0009, Xiaolin Huang, Mengke Li 0001, Yan Yan 0001, Chen Gong 0002, Hanzi Wang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Discrete Elective Hashing with Incomplete Labels for Efficient Cross-Modal RetrievalabstractRecently, supervised cross-modal hashing methods have gained considerable attention due to their ability to mine credible semantic relationships between multi-modal data. These methods typically rely on labels to explore semantic relationships provided that labels are always reliable, which, however, may not be true from the practical perspective. In fact, labels may be incomplete, i.e., true-label incomplete and fine-grained incomplete, which makes the performance of the existing methods deteriorated. To this end, we propose a method called Discrete Elective Hashing with Incomplete Labels (DEH-IL), which is designed to alleviate the impact of incomplete labels. Specifically, we introduce a relaxed label scheme that allows the algorithm to automatically mine potential missing information from incomplete labels, which is beneficial for exploring intra-class relationships. Moreover, we propose a novel elective loss that aggregates all estimations from incomplete labels to mine inter-class relationships. Since elective loss does not rely on any single estimation, it can effectively mitigate estimation errors arising from incomplete labels. By combining these two components, DEH-IL can effectively explore both intra-class and inter-class relationships through incomplete labels. Experimental results on benchmark datasets demonstrate the effectiveness of the proposed method. Donglin Zhang 0001, Changxing Li, Mengke Li 0001, Zhikai Hu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | Feature Fusion from Head to Tail for Long-Tailed Visual RecognitionabstractThe imbalanced distribution of long-tailed data presents a considerable challenge for deep learning models, as it causes them to prioritize the accurate classification of head classes but largely disregard tail classes. The biased decision boundary caused by inadequate semantic information in tail classes is one of the key factors contributing to their low recognition accuracy. To rectify this issue, we propose to augment tail classes by grafting the diverse semantic information from head classes, referred to as head-to-tail fusion (H2T). We replace a portion of feature maps from tail classes with those belonging to head classes. These fused features substantially enhance the diversity of tail classes. Both theoretical analysis and practical experimentation demonstrate that H2T can contribute to a more optimized solution for the decision boundary. We seamlessly integrate H2T in the classifier adjustment stage, making it a plug-and-play module. Its simplicity and ease of implementation allow for smooth integration with existing long-tailed recognition methods, facilitating a further performance boost. Extensive experiments on various long-tailed benchmarks demonstrate the effectiveness of the proposed H2T. The source code is available at https://github.com/Keke921/H2T. Mengke Li 0001, Zhikai Hu, Yang Lu 0009, Weichao Lan, Yiu-Ming Cheung, Hui Huang 0004 |
AAAI | 1 |
| 2024 | Adapt PointFormer: 3D Point Cloud Analysis via Adapting 2D Visual TransformersabstractPre-trained large-scale models have exhibited remarkable efficacy in computer vision, particularly for 2D image analysis. However, when it comes to 3D point clouds, the constrained accessibility of data, in contrast to the vast repositories of images, poses a challenge for the development of 3D pre-trained models. This paper therefore attempts to directly leverage pre-trained models with 2D prior knowledge to accomplish the tasks for 3D point cloud analysis. Accordingly, we propose the Adaptive PointFormer (APF), which fine-tunes pre-trained 2D models with only a modest number of parameters to directly process point clouds, obviating the need for mapping to images. Specifically, we convert raw point clouds into point embeddings for aligning dimensions with image tokens. Given the inherent disorder in point clouds, in contrast to the structured nature of images, we then sequence the point embeddings to optimize the utilization of 2D attention priors. To calibrate attention across 3D and 2D domains and reduce computational overhead, a trainable PointFormer with a limited number of parameters is subsequently concatenated to a frozen pre-trained image model. Extensive experiments on various benchmarks demonstrate the effectiveness of the proposed APF. The source code and more details are available at https://vcc.tech/research/2024/PointFormer. Mengke Li 0001, Yiu-Ming Cheung, Hui Huang 0004 |
ECAI | 1 |
| 2024 | Learning Order Forest for Qualitative-Attribute Data ClusteringabstractClustering is a fundamental approach to understanding data patterns, wherein the intuitive Euclidean distance space is commonly adopted. However, this is not the case for implicit cluster distributions reflected by qualitative attribute values, e.g., the nominal values of attributes like symptoms, marital status, etc. This paper, therefore, discovered a tree-like distance structure to flexibly represent the local order relationship among intra-attribute qualitative values. That is, treating a value as the vertex of the tree allows to capture rich order relationships among the vertex value and the others. To obtain the trees in a clustering-friendly form, a joint learning mechanism is proposed to iteratively obtain more appropriate tree structures and clusters. It turns out that the latent distance space of the whole dataset can be well-represented by a forest consisting of the learned trees. Extensive experiments demonstrate that the joint learning adapts the forest to the clustering task to yield accurate results. Comparisons of 10 counterparts on 12 real benchmark datasets with significance tests verify the superiority of the proposed method. Source code of the proposed method is available at [39]. Mingjie Zhao 0003, Sen Feng, Yiqun Zhang 0006, Mengke Li 0001, Yang Lu 0009, Yiu-Ming Cheung |
ECAI | 4 |
| 2024 | Key Points Centered Sparse Hashing for Cross-Modal RetrievalabstractSupervised cross-modal hashing methods usually construct a massive undirected weighted graph based on labels for training data, with the aim of learning more structured hash codes by preserving relationships within this graph. However, as the volume of data increases, such an approach demands substantial computational and storage resources and tends to aggregate all data points with paths, even semantically unrelated ones, which undermines the retrieval performance. In this paper, we propose to prune less crucial paths from this graph to obtain a clearer representation of relationships among data points. This not only reduces computational resources but separates semantically unrelated data points. Specifically, we define key points within the graph and retain relationships only between all data points and these key points, resulting in a simplified and more transparent graph that is used to supervise hash code learning. Experimental results on three datasets demonstrate that removing unimportant paths from the relationship graph can lead to the learning of more structured hash codes, thereby improving retrieval performance. Zhikai Hu, Yiu-Ming Cheung, Mengke Li 0001, Weichao Lan, Donglin Zhang 0001 |
ICASSP | 3 |
| 2024 | Dynamically Anchored Prompting for Task-Imbalanced Continual Learning
Chenxing Hong, Zhiqi Kang, Mengke Li 0001, Yang Lu 0009, Hanzi Wang |
IJCAI | 5 |
| 2024 | Clustering by Learning the Ordinal Relationships of Qualitative Attribute ValuesabstractIn many real-world clustering tasks, data objects are described by both quantitative and qualitative attributes. Attributes with semantically ordered qualitative values are very common and are usually coded according to their order (i.e., consecutive integers) for clustering. However, semantic order is not always globally interdependent with a certain clustering task. An intuitive case is that level of income (attribute) is not always positively correlated with the level of mental health (label). Using mismatched order surely forms a bottleneck to clustering performance, and conversely, the unsupervised clustering process prevents understanding of "true" order. Therefore, we proposed a novel learning paradigm to tune the value order. More specifically, we adjust the intra-attribute orders, and let this process learn mutually with object clustering, thus bridging the gap between value order and clustering task. To the best of our knowledge, this is the first attempt to learn ordinal relationships among qualitative attribute values. Extensive experiments with significance tests show that our method outperforms the existing relevant clustering approaches on qualitative attribute data. Yiqun Zhang 0006, Yang Lu 0009, Mengke Li 0001, Yiu-Ming Cheung |
IJCNN | 5 |
| 2024 | Improving Visual Prompt Tuning by Gaussian Neighborhood Minimization for Long-Tailed Visual RecognitionabstractLong-tailed visual recognition has received increasing attention recently. Despite fine-tuning techniques represented by visual prompt tuning (VPT) achieving substantial performance improvement by leveraging pre-trained knowledge, models still exhibit unsatisfactory generalization performance on tail classes. To address this issue, we propose a novel optimization strategy called Gaussian neighborhood minimization prompt tuning (GNM-PT), for VPT to address the long-tail learning problem. We introduce a novel Gaussian neighborhood loss, which provides a tight upper bound on the loss function of data distribution, facilitating a flattened loss landscape correlated to improved model generalization. Specifically, GNM-PT seeks the gradient descent direction within a random parameter neighborhood, independent of input samples, during each gradient update. Ultimately, GNM-PT enhances generalization across all classes while simultaneously reducing computational overhead. The proposed GNM-PT achieves state-of-the-art classification accuracies of 90.3%, 76.5%, and 50.1% on benchmark datasets CIFAR100-LT (IR 100), iNaturalist 2018, and Places-LT, respectively. The source code is available at https://github.com/Keke921/GNM-PT. Mengke Li 0001, Yang Lu 0009, Yiqun Zhang 0006, Yiu-Ming Cheung, Hui Huang 0004 |
NeurIPS | 1 |
| 2024 | Audio splicing detection and localization using multistage filterbank spectral sketches and decision fusion
Zhaopin Su, Ziqi Fang, Chensi Lian, Guofu Zhang, Mengke Li 0001 |
Multim. Syst. | 5 |
| 2024 | Cross-Modal Hashing Method With Properties of Hamming Space: A New PerspectiveabstractCross-modal hashing (CMH) has attracted considerable attention in recent years. Almost all existing CMH methods primarily focus on reducing the modality gap and semantic gap, i.e., aligning multi-modal features and their semantics in Hamming space, without taking into account the space gap, i.e., difference between the real number space and the Hamming space. In fact, the space gap can affect the performance of CMH methods. In this paper, we analyze and demonstrate how the space gap affects the existing CMH methods, which therefore raises two problems: solution space compression and loss function oscillation. These two problems eventually cause the retrieval performance deteriorating. Based on these findings, we propose a novel algorithm, namely Semantic Channel Hashing (SCH). First, we classify sample pairs into fully semantic-similar, partially semantic-similar, and semantic-negative ones based on their similarity and impose different constraints on them, respectively, to ensure that the entire Hamming space is utilized. Then, we introduce a semantic channel to alleviate the issue of loss function oscillation. Experimental results on three public datasets demonstrate that SCH outperforms the state-of-the-art methods. Furthermore, experimental validations are provided to substantiate the conjectures regarding solution space compression and loss function oscillation, offering visual evidence of their impact on the CMH methods. Zhikai Hu, Yiu-Ming Cheung, Mengke Li 0001, Weichao Lan |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Compact Neural Network via Stacking Hybrid UnitsabstractAs an effective tool for network compression, pruning techniques have been widely used to reduce the large number of parameters in deep neural networks (NNs). Nevertheless, unstructured pruning has the limitation of dealing with the sparse and irregular weights. By contrast, structured pruning can help eliminate this drawback but it requires complex criteria to determine which components to be pruned. Therefore, this paper presents a new method termed BUnit-Net, which directly constructs compact NNs by stacking designed basic units, without requiring additional judgement criteria anymore. Given the basic units of various architectures, they are combined and stacked systematically to build up compact NNs which involve fewer weight parameters due to the independence among the units. In this way, BUnit-Net can achieve the same compression effect as unstructured pruning while the weight tensors can still remain regular and dense. We formulate BUnit-Net in diverse popular backbones in comparison with the state-of-the-art pruning methods on different benchmark datasets. Moreover, two new metrics are proposed to evaluate the trade-off of compression performance. Experiment results show that BUnit-Net can achieve comparable classification accuracy while saving around 80% FLOPs and 73% parameters. That is, stacking basic units provides a new promising way for network compression. Weichao Lan, Yiu-Ming Cheung, Juyong Jiang, Zhikai Hu, Mengke Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Joint Semantic Preserving Sparse Hashing for Cross-Modal RetrievalabstractSupervised cross-modal hashing has received wide attention in recent years. However, existing methods primarily rely on sample-wise semantic relationships to evaluate the semantic similarity between samples, overlooking the impact of label distribution on enhancing retrieval performance. Moreover, the limited representation capability of traditional dense hash codes hinders the preservation of semantic relationship. To overcome these challenges, we propose a new method, Joint Semantic Preserving Sparse Hashing (JSPSH). Specifically, we introduce a new concept of cluster-wise semantic relationship, which leverages label distribution to indicate which samples are more suitable for clustering. Then, we jointly utilize sample-wise and cluster-wise semantic relationships to supervise the learning of hash codes. In this way, JSPSH preserves both kinds of semantic relationships to ensure that more samples with similar semantics are clustered together, thereby achieving better retrieval results. Furthermore, we utilize high-dimensional sparse hash codes that offer stronger representation capability to preserve such more complex semantics. Finally, an interaction term is introduced in hash functions learning stage to further narrow the gap between modalities. Experimental results on three large-scale datasets demonstrate the effectiveness of JSPSH in achieving superior retrieval performance. Codes are available at https://github.com/hutt94/JSPSH. Zhikai Hu, Yiu-Ming Cheung, Mengke Li 0001, Weichao Lan, Donglin Zhang 0001, Qiang Liu 0018 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Long-Tailed Visual Recognition via Self-Heterogeneous Integration with Knowledge ExcavationabstractDeep neural networks have made huge progress in the last few decades. However, as the real-world data often exhibits a long-tailed distribution, vanilla deep models tend to be heavily biased toward the majority classes. To address this problem, state-of-the-art methods usually adopt a mixture of experts (MoE) to focus on different parts of the long-tailed distribution. Experts in these methods are with the same model depth, which neglects the fact that different classes may have different preferences to be fit by models with different depths. To this end, we propose a novel MoE-based method called Self-Heterogeneous Integration with Knowledge Excavation (SHIKE). We first propose Depth-wise Knowledge Fusion (DKF) to fuse features between different shallow parts and the deep part in one network for each expert, which makes experts more diverse in terms of representation. Based on DKF, we further propose Dynamic Knowledge Transfer (DKT) to reduce the influence of the hardest negative class that has a non-negligible impact on the tail classes in our MoEframework. As a result, the classification accuracy of long-tailed data can be significantly improved, especially for the tail classes. SHIKE achieves the state-of-the-art performance of 56.3%, 60.3%, 75.4% and 41.9% on CIFAR100-LT (IF100), ImageNet-LT, iNaturalist 2018, and Places-LT, respectively. The source code is available at https://github.com/jinyan-06/SHIKE. Mengke Li 0001, Yang Lu 0009, Yiu-Ming Cheung, Hanzi Wang |
CVPR | 2 |
| 2023 | Few-Shot Lip-Password Based Speaker VerificationabstractLip-password has provided a promising solution for speaker verification (Liu and Cheung 2014). Despite the potential of this technology, there are few related studies, largely attributed to the lack of corresponding public datasets. Furthermore, previous works in this field generally demand a substantial amount of training samples and negative samples, impeding their applications from a practical perspective. Therefore, this paper collects a lip-password dataset and proposes a novel few-shot lip-password based speaker verification model, which can be effectively deployed in real-world scenarios because only a small number of data are required for training. Specifically, with an analysis of lip-password features, a down-sampling strategy is presented to generate more training samples. To compensate for the information loss caused by this strategy, a few-shot model, consisting of global and local models, is designed to simultaneously verify the global and local information of the lip-password. Speaker identity is verified only if both stages are passed. The efficacy of the proposed method is demonstrated using the newly collected dataset. Zhikai Hu, Yiu-Ming Cheung, Mengke Li 0001, Weichao Lan |
ICIP | 3 |
| 2023 | DeCAB: Debiased Semi-supervised Learning for Imbalanced Open-Set Data
Xiaolin Huang, Mengke Li 0001, Yang Lu 0009, Hanzi Wang |
PRCV (9) | 2 |
| 2023 | Robust audio copy-move forgery detection on short forged slices using sliding window
Zhaopin Su, Mengke Li 0001, Guofu Zhang, Qinfang Wu, Yaofei Wang |
J. Inf. Secur. Appl. | 2 |
| 2023 | Key Point Sensitive Loss for Long-Tailed Visual RecognitionabstractFor long-tailed distributed data, existing classification models often learn overwhelmingly on the head classes while ignoring the tail classes, resulting in poor generalization capability. To address this problem, we thereby propose a new approach in this paper, in which a key point sensitive (KPS) loss is presented to regularize the key points strongly to improve the generalization performance of the classification model. Meanwhile, in order to improve the performance on tail classes, the proposed KPS loss also assigns relatively large margins on tail classes. Furthermore, we propose a gradient adjustment (GA) optimization strategy to re-balance the gradients of positive and negative samples for each class. By virtue of the gradient analysis of the loss function, it is found that the tail classes always receive negative signals during training, which misleads the tail prediction to be biased towards the head. The proposed GA strategy can circumvent excessive negative signals on tail classes and further improve the overall classification accuracy. Extensive experiments conducted on long-tailed benchmarks show that the proposed method is capable of significantly improving the classification accuracy of the model in tail classes while maintaining competent performance in head classes. Mengke Li 0001, Yiu-Ming Cheung, Zhikai Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | Evolutionary multi-objective optimization for RIS-aided MU-MISO communication systems
Mengke Li 0001, Bai Yan, Jin Zhang 0001 |
Soft Comput. | 1 |
| 2023 | Robust Audio Copy-Move Forgery Detection Using Constant Q Spectral Sketches and GA-SVMabstractAudio recordings used as evidence have become increasingly important to litigation. Before their admissibility as evidence, an audio forensic expert is often required to help determine whether the submitted audio recordings are altered or authentic. Within this field, the copy-move forgery detection (CMFD), which focuses on finding possible forgeries that are derived from the same audio recording, has been an urgent problem in blind audio forensics. However, most of the existing methods require idealistic pre-segmentation and artificial threshold selection to calculate the similarity between segments, which may result in serious misleading and misjudgment especially on high frequency words. In this work, we present a robust method for detecting and locating an audio copy-move forgery on the basis of constant Q spectral sketches (CQSS) and the integration of a customised genetic algorithm (GA) and support vector machine (SVM). Specifically, the CQSS features are first extracted by averaging the logarithm of the squared-magnitude constant Q transform. Then, the CQSS feature set is automatically optimised by a customised GA combined with SVM to obtain the best feature subset and classification model at the same time. Finally, the integrated method, named CQSS-GA-SVM, is evaluated against the state-of-the-art approaches to blind detection of copy-move forgeries on real-world copy-move datasets with read English and Chinese corpus, respectively. The experimental results demonstrate that the proposed CQSS-GA-SVM exhibits significantly high robustness against post-processing based anti-forensics attacks and adaptability to the changes of the duplicated segment duration, the training set size, the recording length, and the forgery type, which may be beneficial to improving the work efficiency of audio forensic experts. Zhaopin Su, Mengke Li 0001, Guofu Zhang, Qinfang Wu, Miqing Li, Weiming Zhang 0001, Xin Yao 0001 |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2022 | Long-tailed Visual Recognition via Gaussian Clouded Logit AdjustmentabstractLong-tailed data is still a big challenge for deep neural networks, even though they have achieved great success on balanced data. We observe that vanilla training on longtailed data with crossentropy loss makes the instance-rich head classes severely squeeze the spatial distribution of the tail classes, which leads to difficulty in classifying tail class samples. Furthermore, the original crossentropy loss can only propagate gradient short-lively because the gradient in softmax form rapidly approaches zero as the logit difference increases. This phenomenon is called softmax saturation. It is unfavorable for training on balanced data, but can be utilized to adjust the validity of the samples in long-tailed data, thereby solving the distorted embedding space of long-tailed problems. To this end, this paper proposes the Gaussian clouded logit adjustment by Gaussian perturbation of different class logits with varied amplitude. We define the amplitude of perturbation as cloud size and set relatively large cloud sizes to tail classes. The large cloud size can reduce the softmax saturation and thereby making tail class samples more active as well as enlarging the embedding space. To alleviate the bias in a classifier, we therefore propose the class-based effective number sampling strategy with classifier retraining. Extensive experiments on benchmark datasets validate the superior performance of the proposed method. Source code is available at https://github.com/Keke921/GCLLoss. Mengke Li 0001, Yiu-Ming Cheung, Yang Lu 0009 |
CVPR | 1 |
| 2022 | Feature-Balanced Loss for Long-Tailed Visual RecognitionabstractDeep neural networks frequently suffer from performance degradation when the training data is long-tailed because several majority classes dominate the training, resulting in a biased model. Recent studies have made a great effort in solving this issue by obtaining good representations from data space, but few of them pay attention to the influence of feature norm on the predicted results. In this paper, we therefore address the long-tailed problem from feature space and thereby propose the feature-balanced loss. Specifically, we encourage larger feature norms of tail classes by giving them relatively stronger stimuli. Moreover, the stimuli intensity is gradually increased in the way of curriculum learning, which improves the generalization of the tail classes, meanwhile maintaining the performance of the head classes. Extensive experiments on multiple popular long-tailed recognition benchmarks demonstrate that the feature-balanced loss achieves superior performance gains compared with the state-of-the-art methods. Mengke Li 0001, Yiu-Ming Cheung, Juyong Jiang |
ICME | 1 |
| 2021 | Facial Structure Guided GAN for Identity-preserved Face Image De-occlusionabstractIn some practical scenarios, such as video surveillance and personal identification, we often have to address the recognition problem of occluded faces, where content replacement by serious occlusion with non-face objects always produces partial appearance and ambiguous representation. Under the circumstances, the performance of face recognition algorithms will often deteriorate to a certain degree. In this paper, we therefore address this problem by removing occlusions on face images and present a new two-stage Facial Structure Guided Generative Adversarial Network (FSG-GAN). In Stage I of the FSG-GAN, the variational auto-encoder is used to predict the facial structure. In Stage II, the predicted facial structure and the occluded image are concatenated and fed into a generative adversarial network (GAN) based model to synthesize the de-occlusion face image. In this way, the facial structure knowledge can be transferred to the synthesis network. Especially, in order to enable the occluded face image to be perceived well, the generator in the GAN based synthesis network utilizes the hybrid dilated convolution modules to extend the receptive field. Furthermore, aiming at further eliminating the appearance ambiguity as well as unnatural texture, a multi-receptive fields discriminator is proposed to utilize the features from different levels. Experiments on the benchmark datasets show the efficacy of the proposed FSG-GAN. Yiu-Ming Cheung, Mengke Li 0001 |
ICMR | 2 |
| 2021 | Iterative Dynamic Generic Learning for Face Recognition From a Contaminated Single-Sample Per PersonabstractThis article focuses on a new and practical problem in single-sample per person face recognition (SSPP FR), i.e., SSPP FR with a contaminated biometric enrolment database (SSPP-ce FR), where the SSPP-based enrolment database is contaminated by nuisance facial variations in the wild, such as poor lightings, expression change, and disguises (e.g., wearing sunglasses, hat, and scarf). In SSPP-ce FR, the most popular generic learning methods will suffer serious performance degradation because the prototype plus variation (P+V) model used in these methods is no longer suitable in such scenarios. The reasons are twofold. First, the contaminated enrolment samples could yield bad prototypes to represent the persons. Second, the generated variation dictionary is simply based on the subtraction of the average face from generic samples of the same person and cannot well depict the intrapersonal variations. To address the SSPP-ce FR problem, we propose a novel iterative dynamic generic learning (IDGL) method, where the labeled enrolment database and the unlabeled query set are fed into a dynamic label feedback network for learning. Specifically, IDGL first recovers the prototypes for the contaminated enrolment samples via a semisupervised low-rank representation (SSLRR) framework and learns a representative variation dictionary by extracting the "sample-specific" corruptions from an auxiliary generic set. Then, it puts them into the P+V model to estimate labels for query samples. Subsequently, the estimated labels will be used as feedback to modify the SSLRR, thus updating new prototypes for the next round of P+V-based label estimation. With the dynamic learning network, the accuracy of the estimated labels is improved iteratively by virtue of the steadily enhanced prototypes. Experiments on various benchmark face data sets have demonstrated the superiority of IDGL over state-of-the-art counterparts. Yiu-Ming Cheung, Qiquan Shi, Mengke Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2020 | Iterative Dynamic Generic Learning For Single Sample Face Recognition With A Contaminated GalleryabstractThis paper studies a new challenging problem in face recognition (FR) with single sample per person (SSPP), i.e., SSPP FR with a contaminated gallery (SSPP-CG FR), where the gallery is contaminated by variations. In SSPP-CG FR, the popular generic learning methods will suffer serious performance degradation because the applied prototype plus variation (P+V) model is not suitable in such scenarios. The reasons are twofold: 1) The contaminated gallery samples yield bad prototypes to represent the persons; 2) The generated variation dictionary is simply based on the subtraction of average face from generic samples of the same person and cannot well depict the intra-personal variations. To tackle SSPPCG FR, we propose a novel Iterative Dynamic Generic Learning (IDGL) method, where the labeled gallery and unlabeled query sets are fed into a dynamic label feedback network for learning. Specifically, IDGL first recovers the prototypes via a semi-supervised low-rank representation (SSLRR) framework and learns a representative variation dictionary by extracting the “sample-specific” corruptions from an auxiliary generic set. Then, it puts them into the P+V model to estimate labels for query samples. Subsequently, the estimated labels are used as the feedbacks to modify the SSLRR, thus updating new prototypes for the next round of P+V based label estimation. With the dynamic learning network, the accuracy of the estimated labels is improved iteratively owing to the steadily enhanced prototypes. Experiments on various benchmark databases have verified the superiority of IDGL. Yiu-Ming Cheung, Qiquan Shi, Mengke Li 0001 |
ICME | 4 |
| 2019 | SAR Image Change Detection Using PCANet Guided by Saliency DetectionabstractThe selection of training samples is important for the accuracy and efficiency of the synthetic aperture radar (SAR) image change detection task. However, training samples are traditionally extracted from the whole image, which leads to longer training time and an unbalanced number of pixels in the changed and unchanged classes. To overcome this problem, we propose a novel change detection method combining saliency detection with a principal component analysis network, named SDPCANet. To enhance the reliability of the training samples and reduce the amount of training samples, the SDPCANet uses context-aware saliency detection to obtain the salient region, from which the training samples are extracted. In addition, to alleviate the gap between the numbers of training samples in two classes, we regulate the candidate samples using the uniform-selecting strategy to enhance the reliability of the training samples for the SDPCANet. Then, the SDPCANet is trained with the extracted training samples and the remaining pixels are classified in the salient region to obtain the final change map. The experimental results on four sets of multitemporal SAR images demonstrate that the SDPCANet outperforms the reference methods proposed recently. Mengke Li 0001, Ming Li 0004, Peng Zhang 0003, Yan Wu 0003, Wanying Song, Lin An |
IEEE Geosci. Remote. Sens. Lett. | 1 |