Jing Li 0009

dblp:l/JingLi9 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
13since 2021 · last 2025
0000-0001-7584-1240ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 2 since 2021Security and privacy · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Uncover and unlearn nuisances: agnostic fully test-time adaptation
Ponhvoan Srey, Yaxin Shi, Hangwei Qian, Jing Li 0009, Ivor W. Tsang
Mach. Learn.4
2025 Alpha and Prejudice: Improving α-Sized Worst Case Fairness via Intrinsic Reweighting
abstract
Achieving worst case group fairness typically relies on maximizing the utility of the worst-off demographic group. However, in practice, demographic information is often unavailable, making direct max-min formulations infeasible. To address this, recent work introduces a relaxed setting, using a lower bound $\alpha $ on the minimal group size-referred to as " $\alpha $ -sized worst case fairness" in this article. We first motivate the importance of this setting by highlighting its relevance to data privacy, a critical yet underexplored perspective. Rather than simply retraining on worst-off samples, we propose a reweighting approach that assigns sample weights based on their intrinsic contributions to fairness. To handle the global nature of worst case objectives efficiently, we develop a stochastic learning algorithm that simplifies training without sacrificing performance. We also address the impact of outliers by introducing a robust variant of our method. Through theoretical analysis and extensive experiments on standard fairness benchmarks, we show that our methods not only connect naturally to existing fairness-through-reweighting approaches but also outperform strong baselines.
Jing Li 0009, Yinghua Yao, Yuangang Pan, Xuanqian Wang, Ivor W. Tsang, Xiuju Fu
IEEE Trans. Neural Networks Learn. Syst.1
2024 Towards Harmless Rawlsian Fairness Regardless of Demographic Prior
abstract
Due to privacy and security concerns, recent advancements in group fairness advocate for model training regardless of demographic information. However, most methods still require prior knowledge of demographics. In this study, we explore the potential for achieving fairness without compromising its utility when no prior demographics are provided to the training set, namely _harmless Rawlsian fairness_. We ascertain that such a fairness requirement with no prior demographic information essential promotes training losses to exhibit a Dirac delta distribution. To this end, we propose a simple but effective method named VFair to minimize the variance of training losses inside the optimal set of empirical losses. This problem is then optimized by a tailored dynamic update approach that operates in both loss and gradient dimensions, directing the model towards relatively fairer solutions while preserving its intact utility. Our experimental findings indicate that regression tasks, which are relatively unexplored from literature, can achieve significant fairness improvement through VFair regardless of any prior, whereas classification tasks usually do not because of their quantized utility measurements. The implementation of our method is publicly available at https://github.com/wxqpxw/VFair.
Xuanqian Wang, Jing Li 0009, Ivor W. Tsang, Yew-Soon Ong
NeurIPS2
2024 Generative Adversarial Ranking Nets
abstract
We propose a new adversarial training framework -- generative adversarial ranking networks (GARNet) to learn from user preferences among a list of samples so as to generate data meeting user-specific criteria. Verbosely, GARNet consists of two modules: a ranker and a generator. The generator fools the ranker to raise generated samples to the top; while the ranker learns to rank generated samples at the bottom. Meanwhile, the ranker learns to rank samples regarding the interested property by training with preferences collected on real samples. The adversarial ranking game between the ranker and the generator enables an alignment between the generated data distribution and the user-preferred data distribution with theoretical guarantees and empirical verification. Specifically, we first prove that when training with full preferences on a discrete property, the learned distribution of GARNet rigorously coincides with the distribution specified by the given score vector based on user preferences. The theoretical results are then extended to partial preferences on a discrete property and further generalized to preferences on a continuous property. Meanwhile, numerous experiments show that GARNet can retrieve the distribution of user-desired data based on full/partial preferences in terms of various interested properties (i.e., discrete/continuous property, single/multiple properties). Code is available at https://github.com/EvaFlower/GARNet.
Yinghua Yao, Yuangang Pan, Jing Li 0009, Ivor W. Tsang, Xin Yao 0001
J. Mach. Learn. Res.3
2024 Sanitized clustering against confounding bias
abstract
Abstract Real-world datasets inevitably contain biases that arise from different sources or conditions during data collection. Consequently, such inconsistency itself acts as a confounding factor that disturbs the cluster analysis. Existing methods eliminate the biases by projecting data onto the orthogonal complement of the subspace expanded by the confounding factor before clustering. Therein, the interested clustering factor and the confounding factor are coarsely considered in the raw feature space, where the correlation between the data and the confounding factor is ideally assumed to be linear for convenient solutions. These approaches are thus limited in scope as the data in real applications is usually complex and non-linearly correlated with the confounding factor. This paper presents a new clustering framework named Sanitized Clustering Against confounding Bias, which removes the confounding factor in the semantic latent space of complex data through a non-linear dependence measure. To be specific, we eliminate the bias information in the latent space by minimizing the mutual information between the confounding factor and the latent representation delivered by variational auto-encoder. Meanwhile, a clustering module is introduced to cluster over the purified latent representations. Extensive experiments on complex datasets demonstrate that our SCAB achieves a significant gain in clustering performance by removing the confounding bias.
Yinghua Yao, Yuangang Pan, Jing Li 0009, Ivor W. Tsang, Xin Yao 0001
Mach. Learn.3
2024 PROUD: PaRetO-gUided diffusion model for multi-objective generation
Yinghua Yao, Yuangang Pan, Jing Li 0009, Ivor W. Tsang, Xin Yao 0001
Mach. Learn.3
2024 Historical Information-Aided Monitoring of Few-Sample Modes in Industrial Processes With Orthogonal Transferred Projection
abstract
Few-sample modes are easy to appear when a new working condition is triggered in industrial processes especially during the early stages of the new working mode. However, monitoring the early behavior of a new mode is important because engineers and operators are less knowledgeable with such a new mode. Considering the few-sample challenge in this problem, a new multisource transfer learning framework is proposed that leverages historical data under various operating conditions to enrich process monitoring over new mode data. In contrast to existing transfer learning-related work, a new unsupervised domain adaptation framework is designed. The historical modes as the source provide precious knowledge and reference to the new mode so that the features of the new mode are robust to noise and insufficient samples. Mathematically, the historical features play the role of a regularizer for the feature learning in the target domain. A geometrical illustration is given and an iterative optimization algorithm is developed with the convergence analysis. Except for the features guided by historical modes, individual features of the new mode are also extracted from the residual part to form a complete monitoring framework. Finally, the effectiveness of the proposed method is validated through a numerical experiment and a real industrial hydrocracking process.
Kai Wang 0024, Xiang Lei, Wenxuan Zhou 0004, Saige Cheng, Jing Li 0009
IEEE Trans. Ind. Informatics5
2024 Taming Overconfident Prediction on Unlabeled Data From Hindsight
abstract
Minimizing prediction uncertainty on unlabeled data is a key factor to achieve good performance in semi-supervised learning (SSL). The prediction uncertainty is typically expressed as the entropy computed by the transformed probabilities in output space. Most existing works distill low-entropy prediction by either accepting the determining class (with the largest probability) as the true label or suppressing subtle predictions (with the smaller probabilities). Unarguably, these distillation strategies are usually heuristic and less informative for model training. From this discernment, this article proposes a dual mechanism, named adaptive sharpening (ADS), which first applies a soft-threshold to adaptively mask out determinate and negligible predictions, and then seamlessly sharpens the informed predictions, distilling certain predictions with the informed ones only. More importantly, we theoretically analyze the traits of ADS by comparing it with various distillation strategies. Numerous experiments verify that ADS significantly improves state-of-the-art SSL methods by making it a plug-in. Our proposed ADS forges a cornerstone for future distillation-based SSL research.
Jing Li 0009, Yuangang Pan, Ivor W. Tsang
IEEE Trans. Neural Networks Learn. Syst.1
2023 Fairness in graph-based semi-supervised learning
abstract
Abstract Machine learning is widely deployed in society, unleashing its power in a wide range of applications owing to the advent of big data. One emerging problem faced by machine learning is the discrimination from data, and such discrimination is reflected in the eventual decisions made by the algorithms. Recent study has proved that increasing the size of training (labeled) data will promote the fairness criteria with model performance being maintained. In this work, we aim to explore a more general case where quantities of unlabeled data are provided, indeed leading to a new form of learning paradigm, namely fair semi-supervised learning. Taking the popularity of graph-based approaches in semi-supervised learning, we study this problem both on conventional label propagation method and graph neural networks, where various fairness criteria can be flexibly integrated. Our developed algorithms are proved to be non-trivial extensions to the existing supervised models with fairness constraints. Extensive experiments on real-world datasets exhibit that our methods achieve a better trade-off between classification accuracy and fairness than the compared baselines.
Tao Zhang 0055, Tianqing Zhu, Mengde Han, Fengwen Chen, Jing Li 0009, Wanlei Zhou 0001, Philip S. Yu
Knowl. Inf. Syst.5
2023 Revisiting model fairness via adversarial examples
Tao Zhang 0055, Tianqing Zhu, Jing Li 0009, Wanlei Zhou 0001, Philip S. Yu
Knowl. Based Syst.3
2023 Earning Extra Performance From Restrictive Feedbacks
abstract
Many machine learning applications encounter situations where model providers are required to further refine the previously trained model so as to gratify the specific need of local users. This problem is reduced to the standard model tuning paradigm if the target data is permissibly fed to the model. However, it is rather difficult in a wide range of practical cases where target data is not shared with model providers but commonly some evaluations about the model are accessible. In this paper, we formally set up a challenge named Earning eXtra PerformancE from restriCTive feEDdbacks (EXPECTED) to describe this form of model tuning problems. Concretely, EXPECTED admits a model provider to access the operational performance of the candidate model multiple times via feedback from a local user (or a group of users). The goal of the model provider is to eventually deliver a satisfactory model to the local user(s) by utilizing the feedbacks. Unlike existing model tuning methods where the target data is always ready for calculating model gradients, the model providers in EXPECTED only see some feedbacks which could be as simple as scalars, such as inference accuracy or usage rate. To enable tuning in this restrictive circumstance, we propose to characterize the geometry of the model performance with regard to model parameters through exploring the parameters' distribution. In particular, for deep models whose parameters distribute across multiple layers, a more query-efficient algorithm is further tailor-designed that conducts layerwise tuning with more attention to those layers which pay off better. Our theoretical analyses justify the proposed algorithms from the aspects of both efficacy and efficiency. Extensive experiments on different applications demonstrate that our work forges a sound solution to the EXPECTED problem, which establishes the foundation for future studies towards this direction.
Jing Li 0009, Yuangang Pan, Yueming Lyu, Yinghua Yao, Yulei Sui, Ivor W. Tsang
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Implicit Weight Learning for Multi-View Clustering
abstract
Exploiting different representations, or views, of the same object for better clustering has become very popular these days, which is conventionally called multi-view clustering. In general, it is essential to measure the importance of each individual view, due to some noises, or inherent capacities in the description. Many previous works model the view importance as weight, which is simple but effective empirically. In this article, instead of following the traditional thoughts, we propose a new weight learning paradigm in the context of multi-view clustering in virtue of the idea of the reweighted approach, and we theoretically analyze its working mechanism. Meanwhile, as a carefully achieved example, all of the views are connected by exploring a unified Laplacian rank constrained graph, which will be a representative method to compare with other weight learning approaches in experiments. Furthermore, the proposed weight learning strategy is much suitable for multi-view data, and it can be naturally integrated with many existing clustering learners. According to the numerical experiments, the proposed implicit weight learning approach is proven effective and practical to use in multi-view clustering.
Feiping Nie 0001, Shaojun Shi, Jing Li 0009, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.3
2022 Fairness in Semi-Supervised Learning: Unlabeled Data Help to Reduce Discrimination
abstract
A growing specter in the rise of machine learning is whether the decisions made by machine learning models are fair. While research is already underway to formalize a machine-learning concept of fairness and to design frameworks for building fair models with sacrifice in accuracy, most are geared toward either supervised or unsupervised learning. Yet two observations inspired us to wonder whether semi-supervised learning might be useful to solve discrimination problems. First, previous study showed that increasing the size of the training set may lead to a better trade-off between fairness and accuracy. Second, the most powerful models today require an enormous of data to train which, in practical terms, is likely possible from a combination of labeled and unlabeled data. Hence, in this paper, we present a framework of fair semi-supervised learning in the pre-processing phase, including pseudo labeling to predict labels for unlabeled data, a re-sampling method to obtain multiple fair datasets and lastly, ensemble learning to improve accuracy and decrease discrimination. A theoretical decomposition analysis of bias, variance and noise highlights the different sources of discrimination and the impact they have on fairness in semi-supervised learning. A set of experiments on real-world and synthetic datasets show that our method is able to use unlabeled data to achieve a better trade-off between accuracy and discrimination.
Tao Zhang 0055, Tianqing Zhu, Jing Li 0009, Mengde Han, Wanlei Zhou 0001, Philip S. Yu
IEEE Trans. Knowl. Data Eng.3
2020 Secure Metric Learning via Differential Pairwise Privacy
abstract
Distance Metric Learning (DML) has drawn much attention over the last two decades. A number of previous works have shown that it performs well in measuring the similarities of individuals given a set of correctly labeled pairwise data by domain experts. These important and precisely-labeled pairwise data are often highly sensitive in real world (e.g., patients similarity). This paper studies, for the first time, how pairwise information can be leaked to attackers during distance metric learning, and develops differential pairwise privacy (DPP), generalizing the definition of standard differential privacy, for secure metric learning. Unlike traditional differential privacy which only applies to independent samples, thus cannot be used for pairwise data, DPP successfully deals with this problem by reformulating the worst case. Specifically, given the pairwise data, we reveal all the involved correlations among pairs in the constructed undirected graph. DPP is then formalized that defines what kind of DML algorithm is private to preserve pairwise data. After that, a case study employing the contrastive loss is exhibited to clarify the details of implementing a DPP- DML algorithm. Particularly, the sensitivity reduction technique is proposed to enhance the utility of the output distance metric. Experiments both on a toy dataset and benchmarks demonstrate that the proposed scheme achieves pairwise data privacy without compromising the output performance much (Accuracy declines less than 0.01 throughout all benchmark datasets when the privacy budget is set at 4).
Jing Li 0009, Yuangang Pan, Yulei Sui, Ivor W. Tsang
IEEE Trans. Inf. Forensics Secur.1
2018 Directly Solving the Original Ratiocut Problem for Effective Data Clustering
abstract
This paper focuses on the original RatioCut problem, which is one of the most representative clustering paradigms. The RatioCut criterion looks for a partition of the graph to achieve the mincut cost while keeping each partition reasonably large. This well-known problem is NP hard and its relaxed form has been widely used in the past several decades. However, the relaxed RatioCut usually suffers two problems: not satisfactory stable clustering performance, and undesired two-stage optimization. In this work, we solve the original RatioCut problem by learning a new similarity matrix which has as many connected components as the cluster number, so that the original RatioCut constraint can be directly satisfied. An easily implemented algorithm is derived to iteratively optimize the proposed method. Experimental results on various real-world benchmark datasets exhibit the effectiveness of the proposed method to solve the RatioCut problem.
Jing Li 0009, Feiping Nie 0001, Xuelong Li 0001
ICASSP1
2018 Auto-Weighted Multi-View Learning for Image Clustering and Semi-Supervised Classification
abstract
Due to the efficiency of learning relationships and complex structures hidden in data, graph-oriented methods have been widely investigated and achieve promising performance. Generally, in the field of multi-view learning, these algorithms construct informative graph for each view, on which the following clustering or classification procedure are based. However, in many real-world data sets, original data always contain noises and outlying entries that result in unreliable and inaccurate graphs, which cannot be ameliorated in the previous methods. In this paper, we propose a novel multi-view learning model which performs clustering/semi-supervised classification and local structure learning simultaneously. The obtained optimal graph can be partitioned into specific clusters directly. Moreover, our model can allocate ideal weight for each view automatically without explicit weight definition and penalty parameters. An efficient algorithm is proposed to optimize this model. Extensive experimental results on different real-world data sets show that the proposed model outperforms other state-of-the-art multi-view algorithms.
Feiping Nie 0001, Guohao Cai, Jing Li 0009, Xuelong Li 0001
IEEE Trans. Image Process.3
2017 Multi-label Learning with Label-Specific Feature Selection
Yan Yan 0025, ShiNing Li, Zhe Yang 0008, Xiao Zhang 0037, Jing Li 0009, Anyi Wang
ICONIP (1)5
2017 Self-weighted Multiview Clustering with Multiple Graphs
abstract
In multiview learning, it is essential to assign a reasonable weight to each view according to its importance. Thus, for multiview clustering task, a wise and elegant method should achieve clustering multiview data while learning the view weights. In this paper, we address this problem by exploring a Laplacian rank constrained graph, which can be approximately as the centroid of the built graph for each view with different confidences. We start our work with a natural thought that the weights can be learned by introducing a hyperparameter. By analyzing the weakness of it, we further propose a new multiview clustering method which is totally self-weighted. Furthermore, once the target graph is obtained in our models, we can directly assign the cluster label to each data point and do not need any postprocessing such as $K$-means in standard spectral clustering. Evaluations on two synthetic datasets prove the effectiveness of our methods. Compared with several representative graph-based multiview clustering approaches on four real-world datasets, experimental results demonstrate that the proposed methods achieve the better performances and our new clustering method is more practical to use.
Feiping Nie 0001, Jing Li 0009, Xuelong Li 0001
IJCAI2
2017 Convex Multiview Semi-Supervised Classification
abstract
In many practical applications, there are a great number of unlabeled samples available, while labeling them is a costly and tedious process. Therefore, how to utilize unlabeled samples to assist digging out potential information about the problem is very important. In this paper, we study a multiclass semi-supervised classification task in the context of multiview data. First, an optimization method named Parametric multiview semi-supervised classification (PMSSC) is proposed, where the built classifier for each individual view is explicitly combined with a weight factor. By analyzing the weakness of it, a new adapted weight learning strategy is further formulated, and we come to the convex multiview semi-supervised classification (CMSSC) method. Comparing with the PMSSC, this method has two significant properties. First, without too much loss in performance, the newly used weight learning technique achieves eliminating a hyperparameter, and thus it becomes more compact in form and practical to use. Second, as its name implies, the CMSSC models a convex problem, which avoids the local-minimum problem. Experimental results on several multiview data sets demonstrate that the proposed methods achieve better performances than recent representative methods and the CMSSC is preferred due to its good traits.
Feiping Nie 0001, Jing Li 0009, Xuelong Li 0001
IEEE Trans. Image Process.2
2016 Parameter-Free Auto-Weighted Multiple Graph Learning: A Framework for Multiview Clustering and Semi-Supervised Classification
Feiping Nie 0001, Jing Li 0009, Xuelong Li 0001
IJCAI2