EDBT 2026 Demo / reviewers in the wild / expert
Weiwei Li 0001
dblp:45/3709-1
· DBLP profile ↗
49ranked-venue papers
8as first author
28since 2021 · last 2026
0000-0001-7811-4719ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 31 · 7 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 since 2021Databases, data management, data science and information retrieval · 7 · 4 since 2021Software engineering, systems software and programming languages · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Learning Label Distribution with Dirichlet Process Mixture ModelabstractLabel Distribution Learning (LDL) is an effective machine learning paradigm for addressing label ambiguity, where each sample is annotated with a distribution that conveys rich semantic information. However, during the actual annotation process of label distributions, annotators often exhibit divergent labeling preferences for the same sample. Most existing LDL methods overlook this heterogeneity, assuming that the observed label distribution originates from a single labeling pattern. Such an assumption limits their capacity to manage inter-annotator disagreement and constrains the generalization of the resulting models. To address this issue, we propose, for the first time, a Dirichlet process mixture model (DPMM)-based framework for LDL. This framework leverages nonparametric Bayesian methods to adaptively uncover diverse latent labeling patterns from the data and to accurately model annotator heterogeneity. Specifically, the ground-truth label distribution of each sample is modeled as a weighted mixture of multiple latent components, where a feature-conditioned gating mechanism adaptively controls the contribution of each component. Experimental results demonstrate that the proposed model consistently achieves competitive performance on several widely-used benchmark datasets. Minglong Wang, Weiwei Li 0001, Yunan Lu 0002, Xiuyi Jia |
AAAI | 2 |
| 2026 | A new open set fault diagnosis method based on adversarial discrimination and deep evidential fusion under limited labeled samples
Weiwei Li 0001, Jinju Zhou, Wei He 0008, You Cao |
Adv. Eng. Informatics | 3 |
| 2025 | Adaptive-Grained Label Distribution LearningabstractLabel polysemy, where an instance can be associated with multiple labels, is common in real-world tasks. LDL (label distribution learning) is an effective learning paradigm for handling label polysemy, where each instance is associated with a label distribution. Although numerous LDL algorithms have been proposed and achieved satisfactory performance on most existing datasets, they are typically trained directly on the collected label distributions which often lack quality guarantees in real-world tasks due to annotator subjectivity and algorithm assumptions. Consequently, direct learning from such uncertain label distributions can lead to unpredictable generalization performance. To address this problem, we propose an adaptive-grained label distribution learning framework whose main idea is to extract relatively reliable supervision information from unreliable label distributions, and thus the label distribution learning task can be decomposed into three subtasks: coarsening label distributions, learning coarse-grained labels and refining coarse-grained labels. In this framework, we design an adaptive label coarsening algorithm to extract an optimal coarsen-grained labels and a label refining function to enhance the coarse-grained label into the final label distributions. Finally, we conduct extensive experiments on real-world datasets to demonstrate the advantages of our proposal. Yunan Lu 0002, Weiwei Li 0001, Dun Liu, Huaxiong Li, Xiuyi Jia |
AAAI | 2 |
| 2025 | CPTCP: Incorporating Convolutional Neural Network into Text-Vector Based Test Case Prioritization for Compilers
Chunqi Li, Weiwei Li 0001, Yaoshen Yu |
ICIC (10) | 2 |
| 2025 | Approximately Correct Label Distribution LearningabstractLabel distribution learning (LDL) is a powerful learning paradigm that emulates label polysemy by assigning label distributions over the label space. However, existing LDL evaluation metrics struggle to capture meaningful performance differences due to their insensitivity to subtle distributional changes, and existing LDL learning objectives often exhibit biases by disproportionately emphasizing a small subset of samples with extreme predictions. As a result, the LDL metrics lose their discriminability, and the LDL objectives are also at risk of overfitting. In this paper, we propose DeltaLDL, a percentage of predictions that are approximately correct within the context of LDL, as a solution to the above problems. DeltaLDL can serve as a novel evaluation metric, which is parameter-free and reflects more on real performance improvements. DeltaLDL can also serve as a novel learning objective, which is differentiable and encourages most samples to be predicted as approximately correct, thereby mitigating overfitting. Our theoretical analysis and empirical results demonstrate the effectiveness of the proposed solution. Weiwei Li 0001, Yunan Lu 0002, Xiuyi Jia |
ICML | 1 |
| 2025 | LIMEFLDL: A Local Interpretable Model-Agnostic Explanations Approach for Label Distribution LearningabstractLabel distribution learning (LDL) is a novel machine learning paradigm that can handle label ambiguity. This paper focuses on the interpretability issue of label distribution learning. Existing local interpretability models are mainly designed for single-label learning problems and are difficult to directly interpret label distribution learning models. In response to this situation, we propose an improved local interpretable model-agnostic explanations algorithm that can effectively interpret any black-box model in label distribution learning. To address the label dependency problem, we introduce the feature attribution distribution matrix and derive the solution formula for explanations under the label distribution form. Meanwhile, to enhance the transparency and trustworthiness of the explanation algorithm, we provide an analytical solution and derive the boundary conditions for explanation convergence and stability. In addition, we design a feature selection scoring function and a fidelity metric for the explanation task of label distribution learning. A series of numerical experiments and human experiments were conducted to validate the performance of the proposed algorithm in practical applications. The experimental results demonstrate that the proposed algorithm achieves high fidelity, consistency, and trustworthiness in explaining LDL models. Xiuyi Jia, Jinchi Li, Yunan Lu 0002, Weiwei Li 0001 |
ICML | 4 |
| 2025 | Divide and Conquer: Learning Label Distribution with SubtasksabstractLabel distribution learning (LDL) is a novel learning paradigm that emulates label polysemy by assigning label distributions over the label space. However, recent LDL work seems to exhibit a notable contradiction: 1) existing LDL methods employ auxiliary tasks to enhance performance, which narrows their focus to specific applications, thereby lacking generalizability; 2) conversely, LDL methods without auxiliary tasks rely on losses tailored solely to the primary task, lacking beneficial data to guide the learning process. In this paper, we propose S-LDL, a novel and minimalist solution that generates subtask label distributions, i.e., a form of extra supervised information, to reconcile the above contradiction. S-LDL encompasses two key aspects: 1) an algorithm capable of generating subtasks without any prior/expert knowledge; and 2) a plug-andplay framework seamlessly compatible with existing LDL methods, and even adaptable to derivative tasks of LDL. Our analysis and experiments demonstrate that S-LDL is effective and efficient. To the best of our knowledge, this paper represents the first endeavor to address LDL via subtasks. Weiwei Li 0001, Xiuyi Jia |
ICML | 2 |
| 2025 | Towards a Pairwise Ranking Model with Orderliness and Monotonicity for Label EnhancementabstractLabel distribution in recent years has been applied in a diverse array of complex decision-making tasks. To address the availability of label distributions, label enhancement has been established as an effective learning paradigm that aims to automatically infer label distributions from readily available multi-label data, e.g., logical labels. Recently, numerous works have demonstrated that the label ranking is significantly beneficial to label enhancement. However, these works still exhibit deficiencies in representing the probabilistic relationships between label distribution and label rankings, or fail to accommodate scenarios where multiple labels are equally important for a given instance. Therefore, we propose PROM, a pairwise ranking model with orderliness and monotonicity, to explain the probabilistic relationship between label distributions and label rankings. Specifically, we propose the monotonicity and orderliness assumptions for the probabilities of different ranking relationships and derive the mass functions for PROM, which are theoretically ensured to preserve the monotonicity and orderliness. Further, we propose a generative label enhancement algorithm based on PROM, which directly learns a label distribution predictor from the readily available multi-label data. Finally, extensive experiments demonstrate the efficacy of our proposed model. Yunan Lu 0002, Yaojin Lin, Weiwei Li 0001, Xiuyi Jia |
NeurIPS | 4 |
| 2025 | CTDip: a diversity-guided test program synthesis approach for boosting compiler bug detection
Junwei Zeng, Weiwei Li 0001 |
Empir. Softw. Eng. | 4 |
| 2025 | Multi-sensor bearing fault diagnosis based on evidential neural network with sensor weights and reliability
Weiwei Li 0001, Wei He 0008, You Cao |
Expert Syst. Appl. | 3 |
| 2025 | Domain Adaptation for Label Distribution LearningabstractLabel distribution learning (LDL) suffers from the dilemma of insufficient target data in real-world applications, while domain adaptation (DA) seems to be able to provide a solution. However, most existing methods of DA, assuming that the instances can correspond to the explicit class information, are devoted only to classification but not to LDL. We argue that indiscriminately applying such DA methods might cause performance degradation in LDL tasks. In this paper, we propose LDL-DA, a novel algorithm dedicated to supervised domain adaptation for label distribution learning, which jointly learns a shared encoding representation from two aspects: 1) contrastive alignment of scarce supervised target data, and 2) minimizing the distance between prototypes of the same label combination. Experiments show that LDL-DA outperforms existing DA methods adapted to LDL, and provides early positive results in DA for LDL. To the best of our knowledge, this paper is the first research on DA for LDL. Weiwei Li 0001, Xiuyi Jia |
IEEE Trans. Big Data | 2 |
| 2024 | Generative Calibration of Inaccurate Annotation for Label Distribution LearningabstractLabel distribution learning (LDL) is an effective learning paradigm for handling label ambiguity. When applying LDL, it typically requires datasets annotated with label distributions. However, obtaining supervised data for LDL is a challenging task. Due to the randomness of label annotation, the annotator can produce inaccurate annotation results for the instance, affecting the accuracy and generalization ability of the LDL model. To address this problem, we propose a generative approach to calibrate the inaccurate annotation for LDL using variational inference techniques. Specifically, we assume that instances with similar features share latent similar label distributions. The feature vectors and label distributions are generated by Gaussian mixture and Dirichlet mixture, respectively. The relationship between them is established through a shared categorical variable, which effectively utilizes the label distribution of instances with similar features, and achieves a more accurate label distribution through the generative approach. Furthermore, we use a confusion matrix to model the factors that contribute to the inaccuracy during the annotation process, which captures the relationship between label distributions and inaccurate label distributions. Finally, the label distribution is used to calibrate the available information in the noisy dataset to obtain the ground-truth label distribution. Yunan Lu 0002, Weiwei Li 0001, Xiuyi Jia |
AAAI | 3 |
| 2024 | CLGNN: UAV Fault Diagnosis via Causal Learning and Graph Neural NetworkabstractTo build an efficient unmanned aerial vehicle (UAV) fault diagnosis system, accurately modeling the relationship between sensor signals is a challenge. A common thought is to use a graph structure for modeling (Graph Neural Network), i.e., sensors as nodes and relationships between sensors are represented by edges. However, the Graph Neural Network (GNN) models, following the "Learning to Attend" principle, aim to maximize the mutual information between features and labels while minimizing training loss, without distinguishing the causal relationships between features and labels. The model’s lack of causal inference capability can lead to instability in model predictions, which ultimately shows up as a degradation in the model’s generalization performance on the test dataset. To address these issues, this paper proposes a GNN with causal learning to implement efficient UAV fault diagnosis, the model is called CLGNN. Specifically, the model first utilizes a GNN to model the sensor signals represented by the graph structure. Then, by intervening at the level of feature extraction, the non-causal components in the features are weakened to improve the model’s generalization performance in fault diagnosis. Extensive experimental results demonstrate that our method can robustly and accurately capture the anomalous signals of UAVs. In addition, we conduct a detailed analysis of the learned graphical representation of the sensor signals, exploring the decision-making basis of the model, which helps to boost the interpretability and reliability of the model. Weiwei Li 0001, Zhuoran Zheng, Xiuyi Jia |
IJCNN | 2 |
| 2024 | Detecting Optimizing Compiler Bugs via History-Driven Test Program MutationabstractCompiler testing is an important task for assuring the quality of compilers. However, most mutation-based compiler testing approaches still suffer from the effectiveness issue due to the ineffective mutation strategies. In this paper, we propose CTHist, which leverages history-driven test program mutation to construct diverse bug-triggering test programs, aiming to discover more optimizing compiler bugs. Specifically, CTHist first examines the test programs that cause historical compiler bugs, and identifies four categories of bug-triggering structures to help generate new bug-triggering test programs. To generate diverse test programs, CTHist then iteratively conducts AST-level mutations to integrate the extracted structures into seed programs, and enhances the connections between the seed and the mutated structures by introducing new control- and data-dependencies. Finally, given the generated test programs, CTHist leverages differentially test based on local hash checksum to test compilers. The experiments on three popular compilers GCC, LLVM and ICX show that, CTHist outperforms four state-of-the-art approaches (i.e., GrayC, Clang-Fuzzer, universalmutator, and Csmith) in detecting more compiler bugs, and improves the Line Coverage of GCC and LLVM by up to 1.15% ∼ 17.77%. Moreover, CTHist has successfully detected 36 new bugs during the practical evaluation, of which 26 are miscompilation bugs (the most difficult to detect), and 27 have been confirmed/fixed by developers. Junwei Zeng, Weiwei Li 0001 |
Internetware | 4 |
| 2024 | ASTSDL: predicting the functionality of incomplete programming code via an AST-sequence-based deep learning model
Yaoshen Yu, Guohua Shen, Weiwei Li 0001, Yichao Shao |
Sci. China Inf. Sci. | 4 |
| 2024 | Sample diversity selection strategy based on label distribution morphology for active label distribution learning
Weiwei Li 0001, Xiuyi Jia |
Pattern Recognit. | 1 |
| 2024 | Adaptive Weighted Ranking-Oriented Label Distribution LearningabstractLabel distribution learning (LDL) is a novel machine-learning paradigm generalized from multilabel learning (MLL). LDL attaches a label distribution to each instance, giving the description degree of different labels. In many real-world applications, key labels, that is, labels with relatively higher description degrees, are preferable to be better predicted. Unfortunately, existing LDL metrics measure the distance or similarity between label distributions from a global perspective, failing to give sufficient attention to key labels. Therefore, we design a novel LDL metric, the description-degree percentile average (DPA), which simultaneously integrates both the exact ranking value and the description degree of each label. The DPA can enhance accuracy in predicting key labels. Furthermore, noting the shape characteristics of the label distributions, we minimize the variance distance between the predicted and the ground-truth label distributions, to better maintain the distinguishability of labels. Finally, we propose an adaptive weighted ranking-oriented LDL algorithm, which is more suitable for realistic LDL problems that require higher accuracy in predicting key labels. We conduct extensive comparison experiments on various types of LDL datasets. Experimental results on both traditional and newly introduced metrics demonstrate the effectiveness of our proposal. Xiuyi Jia, Yunan Lu 0002, Weiwei Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Generative Label Enhancement with Gaussian Mixture and Partial RankingabstractLabel distribution learning (LDL) is an effective learning paradigm for dealing with label ambiguity. When applying LDL, the datasets annotated with label distributions (i.e., the real-valued vectors like the probability distribution) are typically required. Unfortunately, most existing datasets only contain the logical labels, and manual annotating with label distributions is costly. To address this problem, we treat the label distribution as a latent vector and infer its posterior by variational Bayes. Specifically, we propose a generative label enhancement model to encode the process of generating feature vectors and logical label vectors from label distributions in a principled way. In terms of features, we assume that the feature vector is generated by a Gaussian mixture dominated by the label distribution, which captures the one-to-many relationship from the label distribution to the feature vector and thus reduces the feature generation error. In terms of logical labels, we design a probability distribution to generate the logical label vector from a label distribution, which captures partial label ranking in the logical label vector and thus provides a more accurate guidance for inferring the label distribution. Besides, to approximate the posterior of the label distribution, we design a inference model, and derive the variational learning objective. Finally, extensive experiments on real-world datasets validate our proposal. Yunan Lu 0002, Fan Min 0001, Weiwei Li 0001, Xiuyi Jia |
AAAI | 4 |
| 2023 | Label Enhancement via Joint Implicit Representation ClusteringabstractLabel distribution is an effective label form to portray label polysemy (i.e., the cases that an instance can be described by multiple labels simultaneously). However, the expensive annotating cost of label distributions limits its application to a wider range of practical tasks. Therefore, LE (label enhancement) techniques are extensively studied to solve this problem. Existing LE algorithms mostly estimate label distributions by the instance relation or the label relation. However, they suffer from biased instance relations, limited model capabilities, or suboptimal local label correlations. Therefore, in this paper, we propose a deep generative model called JRC to simultaneously learn and cluster the joint implicit representations of both features and labels, which can be used to improve any existing LE algorithm involving the instance relation or local label correlations. Besides, we develop a novel label distribution recovery module, and then integrate it with JRC model, thus constituting a novel generative label enhancement model that utilizes the learned joint implicit representations and instance clusters in a principled way. Finally, extensive experiments validate our proposal. Yunan Lu 0002, Weiwei Li 0001, Xiuyi Jia |
IJCAI | 2 |
| 2023 | Ranking-preserved generative label enhancement
Yunan Lu 0002, Weiwei Li 0001, Huaxiong Li, Xiuyi Jia |
Mach. Learn. | 2 |
| 2023 | Predicting Label Distribution From Tie-Allowed Multi-Label RankingabstractLabel distribution offers more information about label polysemy than logical label. There are presently two approaches to obtaining label distributions: LDL (label distribution learning) and LE (label enhancement). In LDL, experts must annotate training instances with label distributions, and a predictive function is trained on this training set to obtain label distributions. In LE, experts must annotate instances with logical labels, and label distributions are recovered from them. However, LDL is limited by expensive annotations, and LE has no performance guarantee. Therefore, we investigate how to predict label distribution from TMLR (tie-allowed multi-label ranking) which is a compromise on annotation cost but has good performance guarantees. On the one hand, we theoretically dissect the relationship between TMLR and label distribution. We define EAE (expected approximation error) to quantify the quality of an annotation, provide EAE bounds for TMLR, and derive the optimal range of label distributions corresponding to a given TMLR annotation. On the other hand, we propose a framework for predicting label distribution from TMLR via conditional Dirichlet mixtures. This framework blends the procedures of recovering and learning label distributions end-to-end and allows us to effortlessly encode our knowledge by a semi-adaptive scoring function. Extensive experiments validate our proposal. Yunan Lu 0002, Weiwei Li 0001, Huaxiong Li, Xiuyi Jia |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Label Distribution Learning by Maintaining Label Ranking RelationabstractLabel distribution learning (LDL) is a novel machine learning paradigm that can be seen as an extension of multi-label learning (MLL). Compared with MLL, the advantages of LDL are reflected in the following perspectives: (1) the label distribution gives the relevance description of each label to unknown instances in quantitative terms; (2) the distribution implicitly gives the relevance intensities relation of different labels to a particular instance in qualitative terms, i.e., the label ranking relation. All existing LDL models aim to fit the ground-truth label distribution by quantitatively minimizing the distance between distributions or maximizing the similarity between distributions, which only uses the first advantage of the label distribution but ignores the label ranking relation, which may lose some useful semantic information implied in the label distribution, thus reducing the performance of LDL. Therefore, we propose a novel algorithm to solve this problem by introducing the ranking loss function to LDL. In addition, in order to evaluate the LDL algorithms more comprehensively and verify that the ranking loss is beneficial for keeping the label ranking relation, we also introduce two popular ranking evaluation metrics for LDL. The experimental results on 13 real-world datasets validate the effectiveness of our method. Xiuyi Jia, Xiaoxia Shen, Weiwei Li 0001, Yunan Lu 0002, Jihua Zhu |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Filling Missing Labels in Label Distribution Learning by Exploiting Label-Specific Feature SelectionabstractLabel distribution learning (LDL) is a novel learning paradigm for solving label ambiguity problems. The existing LDL algorithms are mainly constructed by considering complete supervised information. However, in practical applications, there are often partial missing labels in the label space, which destroys the structure and relevance between labels and makes it difficult to design accurately learning algorithms. To solve this problem, we propose an incomplete label distribution learning method to fill missing labels by exploiting label-specific feature selection. Firstly, we use the sparse learning to obtain the specific features of each class label, and the common features related to all class labels, respectively. Secondly, we apply low rank constraint to obtain the potential label correlations, and the prediction results are corrected by reconstructing the label distribution consisting of non-missing labels. Finally, a linear mapping relationship between features and labels is constructed to recover missing labels. Experimental results on several data sets demonstrate the effectiveness of the proposed method. Weiwei Li 0001, Yuqing Lu |
IJCNN | 1 |
| 2022 | Label distribution learning with noisy labels via three-way decisions
Weiwei Li 0001, Yuqing Lu, Xiuyi Jia |
Int. J. Approx. Reason. | 1 |
| 2022 | Fast code recommendation via approximate sub-tree matchingabstractSoftware developers often write code that has similar functionality to existing code segments. A code recommendation tool that helps developers reuse these code fragments can significantly improve their efficiency. Several methods have been proposed in recent years. Some use sequence matching algorithms to find the related recommendations. Most of these methods are time-consuming and can leverage only low-level textual information from code. Others extract features from code and obtain similarity using numerical feature vectors. However, the similarity of feature vectors is often not equivalent to the original code’s similarity. Structural information is lost during the process of transforming abstract syntax trees into vectors. We propose an approximate sub-tree matching based method to solve this problem. Unlike existing tree-based approaches that match feature vectors, it retains the tree structure of the query code in the matching process to find code fragments that best match the current query. It uses a fast approximation sub-tree matching algorithm by transforming the sub-tree matching problem into the match between the tree and the list. In this way, the structural information can be used for code recommendation tasks that have high time requirements. We have constructed several real-world code databases covering different languages and granularities to evaluate the effectiveness of our method. The results show that our method outperforms two compared methods, SENSORY and Aroma, in terms of the recall value on all the datasets, and can be applied to large datasets. Yichao Shao, Weiwei Li 0001, Yaoshen Yu |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2022 | ASTENS-BWA: Searching partial syntactic similar regions between source code fragments via AST-based encoded sequence alignment
Yaoshen Yu, Guohua Shen, Weiwei Li 0001, Yichao Shao |
Sci. Comput. Program. | 4 |
| 2021 | Semi-supervised label distribution learning via projection graph embedding
Xiuyi Jia, Weiping Ding 0001, Huaxiong Li, Weiwei Li 0001 |
Inf. Sci. | 5 |
| 2021 | Label Distribution Learning with Label Correlations on Local SamplesabstractLabel distribution learning (LDL) is proposed for solving the label ambiguity problem in recent years, which can be seen as an extension of multi-label learning. To improve the performance of label distribution learning, some existing algorithms exploit label correlations in a global manner that assumes the label correlations are shared by all instances. However, the instances in different groups may share different label correlations, and few label correlations are globally applicable in real-world tasks. In this paper, two novel label distribution learning algorithms are proposed by exploiting label correlations on local samples, which are called GD-LDL-SCL and Adam-LDL-SCL, respectively. To utilize the label correlations on local samples, the influence of local samples is encoded, and a local correlation vector is designed as the additional features for each instance, which is based on the different clustered local samples. Then, the label distribution for an unseen instance can be predicted by exploiting the original features and the additional features simultaneously. Extensive experiments on some real-world data sets validate that our proposed methods can address the label distribution problems effectively and outperform state-of-the-art methods. Xiuyi Jia, Zechao Li, Weiwei Li 0001, Sheng-Jun Huang |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2020 | Tensor-based multi-view label enhancement for multi-label learningabstractLabel enhancement (LE) is a procedure of recovering the label distributions from the logical labels in the multi-label data, the purpose of which is to better represent and mine the label ambiguity problem through the form of label distribution. Existing LE work mainly concentrates on how to leverage the topological information of the feature space and the correlation among the labels, and all are based on single view data. In view of the fact that there are many multi-view data in real-world applications, which can provide richer semantic information from different perspectives, this paper first presents a multi-view label enhancement problem and proposes a tensor-based multi-view label enhancement method, named TMV-LE. Firstly, we introduce the tensor factorization to get the common subspace which contains the high-order relationships among different views. Secondly, we use the common representation and multiple views to jointly mine a more comprehensive topological structure in the dataset. Finally, the topological structure of the feature space is migrated to the label space to get the label distributions. Extensive comparative studies validate that the performance of multi-view multi-label learning can be improved significantly with TMV-LE. Fangwen Zhang, Xiuyi Jia, Weiwei Li 0001 |
IJCAI | 3 |
| 2020 | Privileged label enhancement with multi-label learningabstractLabel distribution learning has attracted more and more attention in view of its more generalized ability to express the label ambiguity. However, it is much more expensive to obtain the label distribution information of the data rather than the logical labels. Thus, label enhancement is proposed to recover the label distributions from the logical labels. In this paper, we propose a novel label enhancement method by using privileged information. We first apply a multi-label learning model to implicitly capture the complex structural information between instances and generate the privileged information. Second, we adopt LUPI (learning with privileged information) paradigm to utilize the privileged information and employ RSVM+ as the prediction model. Finally, comparison experiments on 12 datasets demonstrate that our proposal can better fit the ground-truth label distributions. Wenfang Zhu, Xiuyi Jia, Weiwei Li 0001 |
IJCAI | 3 |
| 2020 | Multi-Label Learning with Local Similarity of SamplesabstractMulti-label learning has been successfully applied to solve instance multi-semantics problems. Moreover, the topology information of samples is often adopted in existing works to improve the prediction performance, in which the similarity of samples is usually calculated in the entire feature space. However, in real-world applications, each label is often determined by a subset of the original features, so when we focus on different labels, the similarity of two instances may be different. In this paper, we propose a multi-label learning method by exploiting the local similarity of samples. Specifically, the smoothness assumption is applied to assume that if the feature subset is similar between samples, the corresponding label should be similar. In addition, L1 regularization is also adopted to sparse the weight coefficients when constraining the output space of the instance. The experimental results on several data sets validate the effectiveness of the proposed method. Wenfang Zhu, Weiwei Li 0001, Xiuyi Jia |
IJCNN | 2 |
| 2020 | ASPDup: AST-Sequence-based Progressive Duplicate Code Detection Tool for Onsite Programming CodeabstractDuplicate code is an example of bad smells, which are usually been refactored after the detection to improve the quality of programs. Locate the duplicate code at the programming phase may reduce the cost of maintenance, but the challenge is it need to detect duplicate code between an incomplete code fragment with complete files, which the existing tools are hard to be applied to this scenario. In this paper, we propose an AST-sequence-based duplicate code detection approach for onsite programming code. The abstract syntax tree (AST) is extracted from source code and then is transformed into an encoded sequence. A local sequence alignment algorithm is used to find highly similar subsequences. After the post-processing, similar regions will be found between two code fragments according to the subsequences. We have developed a prototype tool as a plugin for Visual Studio Code. Experimental results indicate that our approach is effective in finding highly similar regions between cross-granularity code fragments, which can facilitate duplicate code detection for incomplete onsite programming code. Yaoshen Yu, Yu Zhou 0010, Weiwei Li 0001, Yichao Shao |
Internetware | 4 |
| 2020 | Constructing three-way concept lattice based on the composite of classical lattices
Sichun Yang, Yunan Lu 0002, Xiuyi Jia, Weiwei Li 0001 |
Int. J. Approx. Reason. | 4 |
| 2020 | Effort-Aware semi-Supervised just-in-Time defect prediction
Weiwei Li 0001, Wenzhou Zhang, Xiuyi Jia |
Inf. Softw. Technol. | 1 |
| 2020 | Joint Label-Specific Features and Correlation Information for Multi-Label Learning
Xiuyi Jia, Sai-Sai Zhu, Weiwei Li 0001 |
J. Comput. Sci. Technol. | 3 |
| 2019 | SENSORY: Leveraging Code Statement Sequence Information for Code Snippets RecommendationabstractSoftware developers often have to implement unfamiliar programming tasks. When faced with these problems, developers often search online for code snippets as references to learn how to solve the unfamiliar tasks. In recent years, some researchers propose several approaches to use programming context to recommend code snippets. Most of these approaches use information retrieval based techniques and treat code snippets as a set of tokens. However, in code, the smallest meaningful unit is code statement, in general, the line of code. Since these studies did not consider this issue, there is still room for improvement in the code snippets recommendation. In this paper, we propose a code Statement sEquence iNformation baSed cOde snippets Recommendation sYstem (SENSORY). Different from existing token based approaches, SENSORY performs code snippets recommendation at code statement granularity. It uses the Burrows Wheeler Transform algorithm to search relevant code snippets, and uses the structure information to re-rank the results. To evaluate the effectiveness of our proposed method, we construct a code database with 1000000 real world code snippets which contain more than 15000000 lines of code. The experimental results show that SENSORY outperforms the two strong baseline work in terms of precision and NDCG. Lei Ai, Weiwei Li 0001, Yu Zhou 0010, Yaoshen Yu |
COMPSAC (1) | 3 |
| 2019 | Facial Emotion Distribution Learning by Exploiting Low-Rank Label Correlations LocallyabstractEmotion recognition from facial expressions is an interesting and challenging problem and has attracted much attention in recent years. Substantial previous research has only been able to address the ambiguity of “what describes the expression”, which assumes that each facial expression is associated with one or more predefined affective labels while ignoring the fact that multiple emotions always have different intensities in a single picture. Therefore, to depict facial expressions more accurately, this paper adopts a label distribution learning approach for emotion recognition that can address the ambiguity of “how to describe the expression” and proposes an emotion distribution learning method that exploits label correlations locally. Moreover, a local low-rank structure is employed to capture the local label correlations implicitly. Experiments on benchmark facial expression datasets demonstrate that our method can better address the emotion distribution recognition problem than state-of-the-art methods. Xiuyi Jia, Weiwei Li 0001, Changqing Zhang 0002, Zechao Li |
CVPR | 3 |
| 2019 | Label distribution learning with label-specific featuresabstractLabel distribution learning (LDL) is a novel machine learning paradigm to deal with label ambiguity issues by placing more emphasis on how relevant each label is to a particular instance. Many LDL algorithms have been proposed and most of them concentrate on the learning models, while few of them focus on the feature selection problem. All existing LDL models are built on a simple feature space in which all features are shared by all the class labels. However, this kind of traditional data representation strategy tends to select features that are distinguishable for all labels, but ignores label-specific features that are pertinent and discriminative for each class label. In this paper, we propose a novel LDL algorithm by leveraging label-specific features. The common features for all labels and specific features for each label are simultaneously learned to enhance the LDL model. Moreover, we also exploit the label correlations in the proposed LDL model. The experimental results on several real-world data sets validate the effectiveness of our method. Tingting Ren, Xiuyi Jia, Weiwei Li 0001, Zechao Li |
IJCAI | 3 |
| 2019 | Label Distribution Learning with Label Correlations via Low-Rank ApproximationabstractLabel distribution learning (LDL) can be viewed as the generalization of multi-label learning. This novel paradigm focuses on the relative importance of different labels to a particular instance. Most previous LDL methods either ignore the correlation among labels, or only exploit the label correlations in a global way. In this paper, we utilize both the global and local relevance among labels to provide more information for training model and propose a novel label distribution learning algorithm. In particular, a label correlation matrix based on low-rank approximation is applied to capture the global label correlations. In addition, the label correlation among local samples are adopted to modify the label correlation matrix. The experimental results on real-world data sets show that the proposed algorithm outperforms state-of-the-art LDL methods. Tingting Ren, Xiuyi Jia, Weiwei Li 0001 |
IJCAI | 3 |
| 2019 | Effort-Aware Tri-Training for Semi-supervised Just-in-Time Defect Prediction
Wenzhou Zhang, Weiwei Li 0001, Xiuyi Jia |
PAKDD (2) | 2 |
| 2019 | Multi-objective attribute reduction in three-way decision-theoretic rough set model
Weiwei Li 0001, Xiuyi Jia, Bing Zhou 0002 |
Int. J. Approx. Reason. | 1 |
| 2019 | A multiphase cost-sensitive learning method based on the multiclass three-way decision-theoretic rough set model
Xiuyi Jia, Weiwei Li 0001, Lin Shang 0001 |
Inf. Sci. | 2 |
| 2018 | Label Distribution Learning by Exploiting Label CorrelationsabstractLabel distribution learning (LDL) is a newly arisen machine learning method that has been increasingly studied in recent years. In theory, LDL can be seen as a generalization of multi-label learning. Previous studies have shown that LDL is an effective approach to solve the label ambiguity problem. However, the dramatic increase in the number of possible label sets brings a challenge in performance to LDL. In this paper, we propose a novel label distribution learning algorithm to address the above issue. The key idea is to exploit correlations between different labels. We encode the label correlation into a distance to measure the similarity of any two labels. Moreover, we construct a distance-mapping function from the label set to the parameter matrix. Experimental results on eight real label distributed data sets demonstrate that the proposed algorithm performs remarkably better than both the state-of-the-art LDL methods and multi-label learning methods. Xiuyi Jia, Weiwei Li 0001, Junyu Liu, Yu Zhang 0056 |
AAAI | 2 |
| 2018 | Label Distribution Learning by Exploiting Sample Correlations LocallyabstractLabel distribution learning (LDL) is a novel multi-label learning paradigm proposed in recent years for solving label ambiguity. Existing approaches typically exploit label correlations globally to improve the effectiveness of label distribution learning, by assuming that the label correlations are shared by all instances. However, different instances may share different label correlations, and few correlations are globally applicable in real-world applications. In this paper, we propose a new label distribution learning algorithm by exploiting sample correlations locally (LDL-SCL). To encode the influence of local samples, we design a local correlation vector for each instance based on the clustered local samples. Then we predict the label distribution for an unseen instance based on the original features and the local correlation vector simultaneously. Experimental results demonstrate that LDL-SCL can effectively deal with the label distribution problems and perform remarkably better than the state-of-the-art LDL methods. Xiuyi Jia, Weiwei Li 0001 |
AAAI | 3 |
| 2018 | A Novel Takagi-Sugeno Fuzzy System Modeling Method with Joint Feature Selection and Rule ReductionabstractTraditional Takagi-Sugeno (T-S) fuzzy system modeling methods always yield a large number of fuzzy rules. Besides, they also include almost all the original features in the final model. These two factors make the final model sophisticated. In this paper, we propose a novel T-S fuzzy system modeling method called GS-FIS (Group Sparse Fuzzy Inference Systems), which performs fuzzy rule reduction and feature selection simultaneously in a unified framework. Considering the group structure information in the T-S fuzzy system and common features among fuzzy rules, we cast the fuzzy system modeling into a joint group sparse optimization problem and further develop an alternating direction method of multipliers procedure to derive the optimum solution to the problem. Experimental results on the synthetic dataset and several real-world datasets show that the proposed method can not only obtain a satisfactory generalization performance but also reduce the number of fuzzy rules and features effectively. Jun Wang 0024, Jihua Zhu, Yizhang Jiang, Zhaohong Deng, Weiwei Li 0001, Shitong Wang 0001 |
FUZZ-IEEE | 7 |
| 2017 | A Multi-objective Attribute Reduction Method in Decision-Theoretic Rough Set Model
Weiwei Li 0001, Xiuyi Jia, Bing Zhou 0002 |
KSEM | 2 |
| 2017 | Partial order reduction for checking LTL formulae with the next-time operator
Shuanglong Kan, Zhe Chen 0011, Weiwei Li 0001, Yutao Huang |
J. Log. Comput. | 4 |
| 2016 | Neighborhood based decision-theoretic rough set models
Weiwei Li 0001, Xiuyi Jia, Xinye Cai |
Int. J. Approx. Reason. | 1 |
| 2016 | Three-way decisions based software defect prediction
Weiwei Li 0001 |
Knowl. Based Syst. | 1 |