VLDB 2026 Research / reviewers in the wild / expert
Ming Li 0005
dblp:l/MingLi5
· DBLP profile ↗
75ranked-venue papers
9as first author
32since 2021 · last 2026
0000-0001-7977-5500ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 1 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 25 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 18 · 3 first-author · 3 since 2021Software engineering, systems software and programming languages · 10 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ARBench: Algorithmic Reasoner or API Alchemist? Evaluating LLMs Beyond API CallsabstractLarge Language Models (LLMs) have demonstrated impressive capabilities in code generation. Like human programmers, LLMs tend to call high-level APIs and libraries to program efficiently. However, this shortcut may hinder LLMs from learning the essential algorithm reasoning, leading instead to rote memorization of API usage. As a result, LLMs often struggle to generalize to new or domain-specific algorithms that lack ready-made library support. In this work, we propose ARBench, a novel benchmark for evaluating LLMs’ ability to generate machine learning algorithms from scratch, beyond merely invoking high-level APIs. It emphasizes algorithmic reasoning and implementation, distinguishing genuine understanding from superficial API usage. It covers fundamental and advanced machine learning tasks, rigorously assessing current LLMs’ capacity to implement these algorithms from scratch. Our evaluation reveals the strengths and weaknesses of state-of-the-art LLMs in algorithmic reasoning and generalization, offering valuable insights to guide future research and development. Renbiao Liu, Chao-Zeng Ma, Hui Sun 0003, Xin-Ye Li, Ming Li 0005 |
AAAI | 6 |
| 2026 | Dynamic-Static Synergistic Selection Method for Candidate Code Solutions with Generated Test CasesabstractLarge language models (LLMs) show significant improvement in code generation. A common practice is sampling multiple candidate codes to increase the likelihood of producing an accurate solution. However, effectively identifying the best candidate from the pool is a significant challenge. Although existing code consensus methods attempt to solve this issue, they suffer from a critical problem: relying on test cases generated by LLMs, which can be flawed or provide incomplete coverage. This problem can result in erroneous validations, causing correct code to fail flawed tests and preventing the detection of functional differences in candidate code solutions. To address these issues, we present the Dynamic-Static Synergistic Selection Method, a novel framework that combines two complementary analytical approaches. First, it uses the abstract syntax tree (AST) to detect and filter candidate solutions and test cases. Second, the method statically analyzes the quality of the solutions and then dynamically validates functional consistency based on the execution results of the extracted inputs, thereby neutralizing the impact of faulty tests. Extensive experiments demonstrate that this synergistic approach significantly outperforms existing methods, substantially enhancing the correctness of the selected code. Renbiao Liu, Jiang-Tian Xue, Chao-Zeng Ma, Hui Sun 0003, Xin-Ye Li, Ming Li 0005 |
AAAI | 6 |
| 2026 | Mitigating Negative Transfer via Reducing Environmental DisagreementabstractUnsupervised Domain Adaptation (UDA) focuses on transferring knowledge from a labeled source domain to an unlabeled target domain, addressing the challenge of domain shift. Significant domain shifts hinder effective knowledge transfer, leading to negative transfer and deteriorating model performance. Therefore, mitigating negative transfer is essential. This study revisits negative transfer through the lens of causally disentangled learning, emphasizing cross-domain discriminative disagreement on non-causal environmental features as a critical factor. Our theoretical analysis reveals that overreliance on non-causal environmental features as the environment evolves can cause discriminative disagreements (termed environmental disagreement), thereby resulting in negative transfer. To address this, we propose Reducing Environmental Disagreement (RED), which disentangles each sample into domain-invariant causal features and domain-specific non-causal environmental features via adversarially training domain-specific environmental feature extractors in the opposite domains. Subsequently, RED estimates and reduces environmental disagreement based on domain-specific non-causal environmental features. Experimental results confirm that RED effectively mitigates negative transfer and achieves state-of-the-art performance. Hui Sun 0003, Zheng Xie 0001, Haoyuan He 0001, Ming Li 0005 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Mdp3: a Training-Free Approach for List-Wise Frame Selection in Video-Llms
Hui Sun 0003, Shiyin Lu, Weihua Luo, Kaifu Zhang, Ming Li 0005 |
ICCV | 8 |
| 2025 | Revisiting Chain-of-Thought in Code Generation: Do Language Models Need to Learn Reasoning before Coding?abstractLarge Language Models (LLMs) have demonstrated exceptional performance in code generation, becoming increasingly vital for software engineering and development. Recently, Chain-of-Thought (CoT) has proven effective for complex tasks by prompting LLMs to reason step-by-step and provide a final answer.
However, research on *how LLMs learn to reason with CoT data for code generation* remains limited.
In this work, we revisit classic CoT training, which typically learns reasoning steps before the final answer.
We synthesize a dataset to separate the CoT process from code solutions and then conduct extensive experiments to study how CoT works in code generation empirically.
We observe counterintuitive phenomena, suggesting that the traditional training paradigm may not yield benefits for code generation. Instead, training LLMs to generate code first and then output the CoT to explain reasoning steps for code generation is more effective.
Specifically, our results indicate that a 9.86% relative performance improvement can be achieved simply by changing the order between CoT and code. Our findings provide valuable insights into leveraging CoT to enhance the reasoning capabilities of CodeLLMs and improve code generation. Renbiao Liu, Chaoding Yang, Hui Sun 0003, Ming Li 0005 |
ICML | 5 |
| 2025 | APIDocBooster: An Extract-Then-Abstract Framework for Augmenting API DocumentationabstractAPI documentation is often the most trusted resource for programming. Many approaches have been proposed to augment API documentation by summarizing complementary information from external resources like Stack Overflow. Existing extractive summarization approaches excel in producing faithful summaries that accurately represent the source content without input length restrictions. Nevertheless, they suffer from inherent readability limitations. On the other hand, our empirical study on the abstractive-based summarization method, i.e., GPT-4, reveals that GPT-4 can generate coherent and concise summaries but presents limitations in terms of informativeness and faithfulness. We introduce APIDOCBOOSTER, an extract-then-abstract framework that seamlessly fuses the advantages of both extractive (i.e., enabling faithful summaries without length limitation) and abstractive summarization (i.e., producing coherent and concise summaries). APIDocBooster consists of two stages: (1) Context-aware Sentence Section Classification (CSSC) and (2) UPdate SUMmarization (UPSUM). CSSC classifies APIrelevant information collected from multiple sources into API documentation sections. UPSUM generates extractive summaries distinct from original API documentation and then abstractive summaries guided by extractive summaries through in-context learning. To enable automatic evaluation, we construct the first dataset for API documentation augmentation. Our automatic evaluation results reveal that each stage in APIDocBooster outperforms its baselines by a large margin. Our human evaluation also demonstrates the superiority of APIDOCBOOSTER over GPT-4 and shows that it improves the informativeness, relevance and faithfulness by$\mathbf{1 6. 2 2 \%}, \mathbf{1 9. 4 4 \%}$, and$\mathbf{3 7. 1 4 \%}$, respectively. Chengran Yang, Christoph Treude, Yunbo Lyu, Junda He, Ming Li 0005, David Lo 0001 |
ICSME | 7 |
| 2025 | From challenges and pitfalls to recommendations and opportunities: Implementing federated learning in healthcareabstractFederated learning holds great potential for enabling large-scale healthcare research and collaboration across multiple centers while ensuring data privacy and security are not compromised. Although numerous recent studies suggest or utilize federated learning based methods in healthcare, it remains unclear which ones have potential clinical utility. This review paper considers and analyzes the most recent studies up to May 2024 that describe federated learning based methods in healthcare. After a thorough review, we find that the vast majority are not appropriate for clinical use due to their methodological flaws and/or underlying biases which include but are not limited to privacy concerns, generalization issues, and communication costs. As a result, the effectiveness of federated learning in healthcare is significantly compromised. To overcome these challenges, we provide recommendations and promising opportunities that might be implemented to resolve these problems and improve the quality of model development in federated learning with healthcare. • Evaluate recent FL technologies in healthcare, focusing on challenges and pitfalls. • Offer a taxonomic analysis of FL in healthcare across critical aspects. • Recommend strategies for improving FL and ensuring reproducibility. • Highlight trends and opportunities to enhance FL workflow. Ming Li 0005, Zeyu Tang 0001, Guang Yang 0006 |
Medical Image Anal. | 1 |
| 2025 | Probabilistic instance dependent label refinement for noisy label learning
Haoyuan He 0001, Yu Liu 0083, Renbiao Liu, Zheng Xie 0001, Ming Li 0005 |
Mach. Learn. | 5 |
| 2025 | Capturing the context-aware code change via dynamic control flow graph for commit message generation
Yali Du 0002, Yi-Fan Ma, Ming Li 0005 |
Mach. Learn. | 4 |
| 2025 | Deep Rib Fracture Instance Segmentation and Classification From CT on the RibFrac ChallengeabstractRib fractures are a common and potentially severe injury that can be challenging and labor-intensive to detect in CT scans. While there have been efforts to address this field, the lack of large-scale annotated datasets and evaluation benchmarks has hindered the development and validation of deep learning algorithms. To address this issue, the RibFrac Challenge was introduced, providing a benchmark dataset of over 5,000 rib fractures from 660 CT scans, with voxel-level instance mask annotations and diagnosis labels for four clinical categories (buckle, nondisplaced, displaced, or segmental). The challenge includes two tracks: a detection (instance segmentation) track evaluated by an FROC-style metric and a classification track evaluated by an F1-style metric. During the MICCAI 2020 challenge period, 243 results were evaluated, and seven teams were invited to participate in the challenge summary. The analysis revealed that several top rib fracture detection solutions achieved performance comparable or even better than human experts. Nevertheless, the current rib fracture classification solutions are hardly clinically applicable, which can be an interesting area in the future. As an active benchmark and research resource, the data and online evaluation of the RibFrac Challenge are available at the challenge website (https://ribfrac.grand-challenge.org/). In addition, we further analyzed the impact of two post-challenge advancements-large-scale pretraining and rib segmentation-based on our internal baseline for rib fracture detection. These findings lay a foundation for future research and development in AI-assisted rib fracture diagnosis. Jiancheng Yang, Kaiming Kuang, Donglai Wei 0001, Shixuan Gu, Jianying Liu, Zhizhong Chai, Yongjie Xiao, Hao Chen 0011, Liming Xu, Bang Du, Xiangyi Yan, Hao Tang 0010, Adam M. Alessio, Gregory Holste, Jianye He, Lixuan Che, Hanspeter Pfister, Ming Li 0005, Bingbing Ni |
IEEE Trans. Medical Imaging | 24 |
| 2025 | Post-Incorporating Code Structural Knowledge Into Pretrained Models via ICL for Code TranslationabstractCode translation migrates codebases across programming languages. Recently, large language models (LLMs) have achieved significant advancements in software mining. However, handling the syntactic structure of source code remains a challenge. Classic syntax-aware methods depend on intricate model architectures and loss functions, rendering their integration into LLM training resource-intensive. This paper employs in-context learning (ICL), which directly integrates task exemplars into the input context, to post-incorporate code structural knowledge into pre-trained LLMs. We revisit exemplar selection in ICL from an information-theoretic perspective, proposing that list-wise selection based on information coverage is more precise and general objective than traditional methods based on combine similarity and diversity. To address the challenges of quantifying information coverage, we introduce a surrogate measure, Coverage of Abstract Syntax Tree (CAST), measuring maximum subtree coverage between ASTs of test source code and exemplars. Furthermore, we formulate the NP-hard CAST maximization for exemplar selection and prove that it is a standard submodular maximization problem. Therefore, we propose a greedy algorithm for CAST submodular maximization, which theoretically guarantees a (1 − 1/e)-approximate solution in polynomial time complexity. Our method is the first training-free and model-agnostic approach to post-incorporate code structural knowledge into existing LLMs at test time. Experimental results show that our method significantly improves LLMs performance in code translation and reveals two meaningful insights: 1) Code structural knowledge can be effectively post-incorporated into pre-trained LLMs during inference, despite being overlooked during training; 2) Scaling up model size or training data does not lead to the emergence of code structural knowledge, underscoring the necessity of explicitly considering code syntactic structure. Yali Du 0002, Hui Sun 0003, Ming Li 0005 |
IEEE Trans. Software Eng. | 3 |
| 2024 | AUC Optimization from Multiple Unlabeled DatasetsabstractWeakly supervised learning aims to make machine learning more powerful when the perfect supervision is unavailable, and has attracted much attention from researchers. Among the various scenarios of weak supervision, one of the most challenging cases is learning from multiple unlabeled (U) datasets with only a little knowledge of the class priors, or U^m learning for short. In this paper, we study the problem of building an AUC (area under ROC curve) optimal model from multiple unlabeled datasets, which maximizes the pairwise ranking ability of the classifier. We propose U^m-AUC, an AUC optimization approach that converts the U^m data into a multi-label AUC optimization problem, and can be trained efficiently. We show that the proposed U^m-AUC is effective theoretically and empirically. Zheng Xie 0001, Yu Liu 0083, Ming Li 0005 |
AAAI | 3 |
| 2024 | Ambiguity-Aware Abductive LearningabstractAbductive Learning (ABL) is a promising framework for integrating sub-symbolic perception and logical reasoning through abduction. In this case, the abduction process provides supervision for the perception model from the background knowledge. Nevertheless, this process naturally contains uncertainty, since the knowledge base may be satisfied by numerous potential candidates. This implies that the result of the abduction process, i.e., a set of candidates, is ambiguous; both correct and incorrect candidates are mixed in this set. The prior art of abductive learning selects the candidate that has the minimal inconsistency of the knowledge base. However, this method overlooks the ambiguity in the abduction process and is prone to error when it fails to identify the correct candidates. To address this, we propose Ambiguity-Aware Abductive Learning ($\textrm{A}^3\textrm{BL}$), which evaluates all potential candidates and their probabilities, thus preventing the model from falling into sub-optimal solutions. Both experimental results and theoretical analyses prove that $\textrm{A}^3\textrm{BL}$ markedly enhances ABL by efficiently exploiting the ambiguous abduced supervision. Haoyuan He 0001, Hui Sun 0003, Zheng Xie 0001, Ming Li 0005 |
ICML | 4 |
| 2024 | Efficient and Stable Offline-to-online Reinforcement Learning via Continual Policy Revitalization
Chenyang Wu 0001, Chenxiao Gao, Zongzhang Zhang, Ming Li 0005 |
IJCAI | 5 |
| 2024 | A Joint Learning Model with Variational Interaction for Multilingual Program TranslationabstractPrograms implemented in various programming languages form the foundation of software applications. To alleviate the burden of program migration and facilitate the development of software systems, automated program translation across languages has garnered significant attention. Previous approaches primarily focus on pairwise translation paradigms, learning translation between pairs of languages using bilingual parallel data. However, parallel data is difficult to collect for some language pairs, and the distribution of program semantics across languages can shift, posing challenges for pairwise program translation. In this paper, we argue that jointly learning a unified model to translate code across multiple programming languages is superior to separately learning from bilingual parallel data. We propose Variational Interaction for Multilingual Program Translation (VIM-PT), a disentanglement-based generative approach that jointly trains a unified model for multilingual program translation across multiple languages. VIM-PT disentangles code into language-shared and language-specific features, using variational inference and interaction information with a novel lower bound, then achieves program translation through conditional generation. VIM-PT demonstrates four advantages: 1) captures language-shared information more accurately from various implementations and improves the quality of multilingual program translation, 2) mines and leverages the capability of non-parallel data, 3) addresses the distribution shift of program semantics across languages, 4) and serves as a unified model, reducing deployment complexity. Yali Du 0002, Hui Sun 0003, Ming Li 0005 |
ASE | 3 |
| 2024 | Ranking-Aware Unbiased Post-Click Conversion Rate Estimation via AUC Optimization on Entire Exposure SpaceabstractEstimating the post-click conversion rate (CVR) accurately in ranking systems is crucial in industrial applications. However, this task is often challenged by data sparsity and selection bias, which hinder accurate ranking. Previous approaches to address these challenges have typically focused on either modeling CVR across the entire exposure space which includes all exposure events, or providing unbiased CVR estimation separately. However, the lack of integration between these objectives has limited the overall performance of CVR estimation. Therefore, there is a pressing need for a method that can simultaneously provide unbiased CVR estimates across the entire exposure space. To achieve it, we formulate the CVR estimation task as an Area Under the Curve (AUC) optimization problem and propose the Entire-space Weighted AUC (EWAUC) framework. EWAUC utilizes sample reweighting techniques to handle selection bias and employs pairwise AUC risk, which incorporates more information from limited clicked data, to handle data sparsity. In order to model CVR across the entire exposure space unbiasedly, EWAUC treats the exposure data as both conversion data and non-conversion data to calculate the loss. The properties of AUC risk guarantee the unbiased nature of the entire space modeling. We provide comprehensive theoretical analysis to validate the unbiased nature of our approach. Additionally, extensive experiments conducted on real-world datasets demonstrate that our approach outperforms state-of-the-art methods in terms of ranking performance for the CVR estimation task. Yu Liu 0083, Qinglin Jia, Chuhan Wu, Zhaocheng Du, Zheng Xie 0001, Ruiming Tang, Muyu Zhang, Ming Li 0005 |
RecSys | 9 |
| 2024 | Reduced implication-bias logic loss for neuro-symbolic learning
Haoyuan He 0001, Wang-Zhou Dai, Ming Li 0005 |
Mach. Learn. | 3 |
| 2024 | Weakly Supervised AUC Optimization: A Unified Partial AUC ApproachabstractSince acquiring perfect supervision is usually difficult, real-world machine learning tasks often confront inaccurate, incomplete, or inexact supervision, collectively referred to as weak supervision. In this work, we present WSAUC, a unified framework for weakly supervised AUC optimization problems, which covers noisy label learning, positive-unlabeled learning, multi-instance learning, and semi-supervised learning scenarios. Within the WSAUC framework, we first frame the AUC optimization problems in various weakly supervised scenarios as a common formulation of minimizing the AUC risk on contaminated sets, and demonstrate that the empirical risk minimization problems are consistent with the true AUC. Then, we introduce a new type of partial AUC, specifically, the reversed partial AUC (rpAUC), which serves as a robust training objective for AUC maximization in the presence of contaminated labels. WSAUC offers a universal solution for AUC optimization in various weakly supervised scenarios by maximizing the empirical rpAUC. Theoretical and experimental results under multiple settings support the effectiveness of WSAUC on a range of weakly supervised AUC optimization tasks. Zheng Xie 0001, Yu Liu 0083, Haoyuan He 0001, Ming Li 0005, Zhi-Hua Zhou |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | : A Large-Scale Benchmark for Rib Labeling and Anatomical Centerline ExtractionabstractAutomatic rib labeling and anatomical centerline extraction are common prerequisites for various clinical applications. Prior studies either use in-house datasets that are inaccessible to communities, or focus on rib segmentation that neglects the clinical significance of rib labeling. To address these issues, we extend our prior dataset (RibSeg) on the binary rib segmentation task to a comprehensive benchmark, named RibSeg v2, with 660 CT scans (15,466 individual ribs in total) and annotations manually inspected by experts for rib labeling and anatomical centerline extraction. Based on the RibSeg v2, we develop a pipeline including deep learning-based methods for rib labeling, and a skeletonization-based method for centerline extraction. To improve computational efficiency, we propose a sparse point cloud representation of CT scans and compare it with standard dense voxel grids. Moreover, we design and analyze evaluation metrics to address the key challenges of each task. Our dataset, code, and model are available online to facilitate open research at https://github.com/M3DV/RibSeg. Shixuan Gu, Donglai Wei 0001, Jason Ken Adhinarta, Kaiming Kuang, Yongjie Jessica Zhang, Hanspeter Pfister, Bingbing Ni, Jiancheng Yang, Ming Li 0005 |
IEEE Trans. Medical Imaging | 10 |
| 2024 | GMILT: A Novel Transformer Network That Can Noninvasively Predict EGFR Mutation StatusabstractNoninvasively and accurately predicting the epidermal growth factor receptor (EGFR) mutation status is a clinically vital problem. Moreover, further identifying the most suspicious area related to the EGFR mutation status can guide the biopsy to avoid false negatives. Deep learning methods based on computed tomography (CT) images may improve the noninvasive prediction of EGFR mutation status and potentially help clinicians guide biopsies by visual methods. Inspired by the potential inherent links between EGFR mutation status and invasiveness information, we hypothesized that the predictive performance of a deep learning network can be improved through extra utilization of the invasiveness information. Here, we created a novel explainable transformer network for EGFR classification named gated multiple instance learning transformer (GMILT) by integrating multi-instance learning and discriminative weakly supervised feature learning. Pathological invasiveness information was first introduced into the multitask model as embeddings. GMILT was trained and validated on a total of 512 patients with adenocarcinoma and tested on three datasets (the internal test dataset, the external test dataset, and The Cancer Imaging Archive (TCIA) public dataset). The performance (area under the curve (AUC) =0.772 on the internal test dataset) of GMILT exceeded that of previously published methods and radiomics-based methods (i.e., random forest and support vector machine) and attained a preferable generalization ability (AUC =0.856 in the TCIA test dataset and AUC =0.756 in the external dataset). A diameter-based subgroup analysis further verified the efficiency of our model (most of the AUCs exceeded 0.772) to noninvasively predict EGFR mutation status from computed tomography (CT) images. In addition, because our method also identified the "core area" of the most suspicious area related to the EGFR mutation status, it has the potential ability to guide biopsies. Wei Zhao 0040, Weidao Chen, Du Lei, Jiancheng Yang, Yanjing Chen, Yingjia Jiang, Jiangfen Wu, Bingbing Ni, Yeqi Sun, Yingli Sun, Ming Li 0005, Jun Liu 0075 |
IEEE Trans. Neural Networks Learn. Syst. | 13 |
| 2023 | Cooperative and Adversarial Learning: Co-enhancing Discriminability and Transferability in Domain AdaptationabstractDiscriminability and transferability are two goals of feature learning for domain adaptation (DA), as we aim to find the transferable features from the source domain that are helpful for discriminating the class label in the target domain. Modern DA approaches optimize discriminability and transferability by adopting two separate modules for the two goals upon a feature extractor, but lack fully exploiting their relationship. This paper argues that by letting the discriminative module and transfer module help each other, better DA can be achieved. We propose Cooperative and Adversarial LEarning (CALE) to combine the optimization of discriminability and transferability into a whole, provide one solution for making the discriminative module and transfer module guide each other. Specifically, CALE generates cooperative (easy) examples and adversarial (hard) examples with both discriminative module and transfer module. While the easy examples that contain the module knowledge can be used to enhance each other, the hard ones are used to enhance the robustness of the corresponding goal. Experimental results show the effectiveness of CALE for unifying the learning of discriminability and transferability, as well as its superior performance. Hui Sun 0003, Zheng Xie 0001, Xin-Ye Li, Ming Li 0005 |
AAAI | 4 |
| 2023 | Semi-supervised Learning with Support Isolation by Small-Paced Self-TrainingabstractIn this paper, we address a special scenario of semi-supervised learning, where the label missing is caused by a preceding filtering mechanism, i.e., an instance can enter a subsequent process in which its label is revealed if and only if it passes the filtering mechanism. The rejected instances are prohibited to enter the subsequent labeling process due to economical or ethical reasons, making the support of the labeled and unlabeled distributions isolated from each other. In this case, semi-supervised learning approaches which rely on certain coherence of the labeled and unlabeled distribution would suffer from the consequent distribution mismatch, and hence result in poor prediction performance. In this paper, we propose a Small-Paced Self-Training framework, which iteratively discovers labeled and unlabeled instance subspaces with bounded Wasserstein distance. We theoretically prove that such a framework may achieve provably low error on the pseudo labels during learning. Experiments on both benchmark and pneumonia diagnosis tasks show that our method is effective. Zheng Xie 0001, Hui Sun 0003, Ming Li 0005 |
AAAI | 3 |
| 2023 | Beyond Lexical Consistency: Preserving Semantic Consistency for Program TranslationabstractProgram translation aims to convert the input programs from one programming language to another. Automatic program translation is a prized target of software engineering research, which leverages the reusability of projects and improves the efficiency of development. Recently, thanks to the rapid development of deep learning model architectures and the availability of large-scale parallel corpus of programs, the performance of program translation has been greatly improved. However, the existing program translation models are still far from satisfactory, in terms of the quality of translated programs. In this paper, we argue that a major limitation of the current approaches is the lack of consideration of semantic consistency. Beyond lexical consistency, semantic consistency is also critical for the task. To make the program translation model more semantically aware, we propose a general framework named Preserving Semantic Consistency for Program Translation (PSCPT), which considers semantic consistency with regularization in the training objective of program translation and can be easily applied to all encoder-decoder methods with various neural networks (e.g., LSTM, Transformer) as the backbone. We conduct extensive experiments in 7 general programming languages. Experimental results show that with CodeBERT as the backbone, our approach outperforms not only the state-of-the-art open-source models but also the commercial closed large language models (e.g., textdavinci-002, text-davinci-003) on the program translation task. Our replication package (including code, data, etc.) is publicly available at https://github.com/duyali2000/PSCPT. Yali Du 0002, Yi-Fan Ma, Zheng Xie 0001, Ming Li 0005 |
ICDM | 4 |
| 2023 | CHRONOS: Time-Aware Zero-Shot Identification of Libraries from Vulnerability ReportsabstractTools that alert developers about library vulnerabilities depend on accurate, up-to-date vulnerability databases which are maintained by security researchers. These databases record the libraries related to each vulnerability. However, the vulnerability reports may not explicitly list every library and human analysis is required to determine all the relevant libraries. Human analysis may be slow and expensive, which motivates the need for automated approaches. Researchers and practitioners have proposed to automatically identify libraries from vulnerability reports using extreme multi-label learning (XML). While state-of-the-art XML techniques showed promising performance, their experimental settings do not practically fit what happens in reality. Previous studies randomly split the vulnerability reports data for training and testing their models without considering the chronological order of the reports. This may unduly train the models on chronologically newer reports while testing the models on chronologically older ones. However, in practice, one often receives chronologically new reports, which may be related to previously unseen libraries. Under this practical setting, we observe that the performance of current XML techniques declines substantially, e.g., F1 decreased from 0.7 to 0.24 under experiments without and with consideration of chronological order of vulnerability reports. We propose a practical library identification approach, namely Chronos, based on zero-shot learning. The novelty of Chronos is three-fold. First, Chronos fits into the practical pipeline by considering the chronological order of vulnerability reports. Second, Chronos enriches the data of the vulnerability descriptions and labels using a carefully designed data enhancement step. Third, Chronos exploits the temporal ordering of the vulnerability reports using a cache to prioritize prediction of versions of libraries that recently had reports of vulnerabilities. In our experiments, Chronos achieves an average F1-score of 0.75, 3x better than the best XML-based approach. Data enhancement and the time-aware adjustment improve Chronos over the vanilla zero-shot learning model by 27% in average F1. Yunbo Lyu, Thanh Le-Cong, Hong Jin Kang, Ratnadira Widyasari, Bach Le 0001, Ming Li 0005, David Lo 0001 |
ICSE | 7 |
| 2023 | Capturing the Long-Distance Dependency in the Control Flow Graph via Structural-Guided Attention for Bug LocalizationabstractTo alleviate the burden of software maintenance, bug localization, which aims to automatically locate the buggy source files based on the bug report, has drawn significant attention in the software mining community. Recent studies indicate that the program structure in source code carries more semantics reflecting the program behavior, which is beneficial for bug localization. Benefiting from the rich structural information in the Control Flow Graph (CFG), CFG-based bug localization methods have achieved the state-of-the-art performance. Existing CFG-based methods extract the semantic feature from the CFG via the graph neural network. However, the step-wise feature propagation in the graph neural network suffers from the problem of information loss when the propagation distance is long, while the long-distance dependency is rather common in the CFG. In this paper, we argue that the long-distance dependency is crucial for feature extraction from the CFG, and propose a novel bug localization model named sgAttention. In sgAttention, a particularly designed structural-guided attention is employed to globally capture the information in the CFG, where features of irrelevant nodes are masked for each node to facilitate better feature extraction from the CFG. Experimental results on four widely-used open-source software projects indicate that sgAttention averagely improves the state-of-the-art bug localization methods by 32.9\% and 29.2\% and the state-of-the-art pre-trained models by 5.8\% and 4.9\% in terms of MAP and MRR, respectively. Yi-Fan Ma, Yali Du 0002, Ming Li 0005 |
IJCAI | 3 |
| 2023 | Enhancing unsupervised domain adaptation by exploiting the conceptual consistency of multiple self-supervised tasks
Hui Sun 0003, Ming Li 0005 |
Sci. China Inf. Sci. | 2 |
| 2023 | Machine/Deep Learning for Software Engineering: A Systematic Literature ReviewabstractSince 2009, the deep learning revolution, which was triggered by the introduction of ImageNet, has stimulated the synergy between Software Engineering (SE) and Machine Learning (ML)/Deep Learning (DL). Meanwhile, critical reviews have emerged that suggest that ML/DL should be used cautiously. To improve the applicability and generalizability of ML/DL-related SE studies, we conducted a 12-year Systematic Literature Review (SLR) on 1,428 ML/DL-related SE papers published between 2009 and 2020. Our trend analysis demonstrated the impacts that ML/DL brought to SE. We examined the complexity of applying ML/DL solutions to SE problems and how such complexity led to issues concerning the reproducibility and replicability of ML/DL studies in SE. Specifically, we investigated how ML and DL differ in data preprocessing, model training, and evaluation when applied to SE tasks, and what details need to be provided to ensure that a study can be reproduced or replicated. By categorizing the rationales behind the selection of ML/DL techniques into five themes, we analyzed how model performance, robustness, interpretability, complexity, and data simplicity affected the choices of ML/DL models. LiGuo Huang, Amiao Gao, Jidong Ge, Haitao Feng, Ishna Satyarth, Ming Li 0005, He Zhang 0001, Vincent Ng 0001 |
IEEE Trans. Software Eng. | 8 |
| 2022 | Learning from the Multi-Level Abstraction of the Control Flow Graph via Alternating Propagation for Bug LocalizationabstractBug localization aims to automatically locate the buggy source files according to the bug report, which plays an important role in software maintenance. Recent studies indicate that exploiting the program structure is beneficial for bug localization. Benefiting from the rich statement-level implementation detail in the Control Flow Graph (CFG), CFG-based bug localization methods have achieved state-of-the-art performance. However, due to the huge semantic gap between the high-level description of the unexpected program behavior in the bug report and the low-level implementation detail in the CFG, it is challenging to directly establish the match between the bug report and the CFG. In this paper, we argue that this gap can be bridged through the multi-level abstraction of the CFG, where each node in an abstraction level corresponds to a code block with a certain granularity. The multi-level abstraction of the CFG reflects the essence of the structured programming paradigm. We further propose a novel model named MLA (Multi-Level Abstraction of the control flow graph) for bug localization, which contains a particularly designed model that alternately propagates the block feature within and between abstraction levels, corresponding to reflecting the control flow and summarizing the block functionality, respectively. Experimental results on four widely-used open-source software projects show that MLA outperforms the state-of-the-art bug localization methods. Yi-Fan Ma, Ming Li 0005 |
ICDM | 2 |
| 2022 | The flowing nature matters: feature learning from the control flow graph of source code for bug localization
Yi-Fan Ma, Ming Li 0005 |
Mach. Learn. | 2 |
| 2021 | Enhancing Context-Based Meta-Reinforcement Learning Algorithms via An Efficient Task Encoder (Student Abstract)abstractMeta-Reinforcement Learning (meta-RL) algorithms enable agents to adapt to new tasks from small amounts of exploration, based on the experience of similar tasks. Recent studies have pointed out that a good representation of a task is key to the success of off-policy context-based meta-RL. Inspired by contrastive methods in unsupervised representation learning, we propose a new method to learn the task representation based on the mutual information between transition tuples in a trajectory and the task embedding. We also propose a new estimation for task similarity based on Q-function, which can be used to form a constraint on the distribution of the encoded task variables, making the task encoder encode the task variables more effective on new tasks. Experiments on meta-RL tasks show that the newly proposed method outperforms existing meta-RL algorithms. Feng Xu 0007, Shengyi Jiang, Zongzhang Zhang, Yang Yu 0001, Ming Li 0005, Dong Li 0016, Wulong Liu |
AAAI | 6 |
| 2021 | Towards Generating Summaries for Lexically Confusing Code through Code ErosionabstractCode summarization aims to summarize code functionality as high-level nature language descriptions to assist in code comprehension. Recent approaches in this field mainly focus on generating summaries for code with precise identifier names, in which meaningful words can be found indicating code functionality. When faced with lexically confusing code, current approaches are likely to fail since the correlation between code lexical tokens and summaries is scarce. To tackle this problem, we propose a novel summarization framework named VECOS. VECOS introduces an erosion mechanism to conquer the model's reliance on precisely defined lexical information. To facilitate learning the eroded code's functionality, we force the representation of the eroded code to align with the representation of its original counterpart via variational inference. Experimental results show that our approach outperforms the state-of-the-art approaches to generate coherent and reliable summaries for various lexically confusing code. Fan Yan, Ming Li 0005 |
IJCAI | 2 |
| 2021 | Deep Transfer Bug LocalizationabstractMany projects often receive more bug reports than what they can handle. To help debug and close bug reports, a number of bug localization techniques have been proposed. These techniques analyze a bug report and return a ranked list of potentially buggy source code files. Recent development on bug localization has resulted in the construction of effective supervised approaches that use historical data of manually localized bugs to boost performance. Unfortunately, as highlighted by Zimmermann et al., sufficient bug data is often unavailable for many projects and companies. This raises the need for cross-project bug localization - the use of data from a project to help locate bugs in another project. To fill this need, we propose a deep transfer learning approach for cross-project bug localization. Our proposed approach named TRANP-CNN extracts transferable semantic features from source project and fully exploits labeled data from target project for effective cross-project bug localization. We have evaluated TRANP-CNN on curated high-quality bug datasets and our experimental results show that TRANP-CNN can locate buggy files correctly at top 1, top 5, and top 10 positions for 29.9, 51.7, 61.3 percent of the bugs respectively, which significantly outperform state-of-the-art bug localization solution based on deep learning and several other advanced alternative solutions considering various standard evaluation metrics. Xuan Huo, Ferdian Thung, Ming Li 0005, David Lo 0001 |
IEEE Trans. Software Eng. | 3 |
| 2020 | Control Flow Graph Embedding Based on Multi-Instance Decomposition for Bug LocalizationabstractDuring software maintenance, bug report is an effective way to identify potential bugs hidden in a software system. It is a great challenge to automatically locate the potential buggy source code according to a bug report. Traditional approaches usually represent bug reports and source code from a lexical perspective to measure their similarities. Recently, some deep learning models are proposed to learn the unified features by exploiting the local and sequential nature, which overcomes the difficulty in modeling the difference between natural and programming languages. However, only considering local and sequential information from one dimension is not enough to represent the semantics, some multi-dimension information such as structural and functional nature that carries additional semantics has not been well-captured. Such information beyond the lexical and structural terms is extremely vital in modeling program functionalities and behaviors, leading to a better representation for identifying buggy source code. In this paper, we propose a novel model named CG-CNN, which is a multi-instance learning framework that enhances the unified features for bug localization by exploiting structural and sequential nature from the control flow graph. Experimental results on widely-used software projects demonstrate the effectiveness of our proposed CG-CNN model. Xuan Huo, Ming Li 0005, Zhi-Hua Zhou |
AAAI | 2 |
| 2020 | Deep Time-Stream Framework for Click-through Rate Prediction by Tracking Interest EvolutionabstractClick-through rate (CTR) prediction is an essential task in industrial applications such as video recommendation. Recently, deep learning models have been proposed to learn the representation of users' overall interests, while ignoring the fact that interests may dynamically change over time. We argue that it is necessary to consider the continuous-time information in CTR models to track user interest trend from rich historical behaviors. In this paper, we propose a novel Deep Time-Stream framework (DTS) which introduces the time information by an ordinary differential equations (ODE). DTS continuously models the evolution of interests using a neural network, and thus is able to tackle the challenge of dynamically representing users' interests based on their historical behaviors. In addition, our framework can be seamlessly applied to any existing deep CTR models by leveraging the additional Time-Stream Module, while no changes are made to the original CTR models. Experiments on public dataset as well as real industry dataset with billions of samples demonstrate the effectiveness of proposed approaches, which achieve superior performance compared with existing methods. Wenhao Zheng 0001, Yao Hu 0002, Jianke Zhu, Ming Li 0005 |
AAAI | 7 |
| 2020 | Learning Code Changes by Exploiting Bidirectional Converting DeviationabstractSoftware systems evolve with constant code changes when requirements change or bugs are found. Assessing the quality of code change is a vital part of software development. However, most existing software mining methods inspect software data from a static view and learn global code semantics from a snapshot of code, which cannot capture the semantic information of small changes and are under{-}representation for rich historical code changes. How to build a model to emphasize the code change remains a great challenge. In this paper, we propose a novel deep neural network called CCL, which models a forward converting process from the code before change to the code after change and a backward converting process inversely, and the change representations of bidirectional converting processes can be learned. By exploiting the deviation of the converting processes, the code change can be evaluated by the network. Experimental results on open source projects indicate that CCL significantly outperforms the compared methods in code change learning. Jia-Wei Mi, Ming Li 0005 |
ACML | 3 |
| 2020 | Automatic fetal brain extraction from 2D in utero fetal MRI slices using deep neural network
Yishan Luo, Lin Shi 0001, Xin Zhang 0013, Ming Li 0005, Bing Zhang 0012, Defeng Wang |
Neurocomputing | 5 |
| 2020 | Enhancing supervised bug localization with metadata and stack-trace
Yaojing Wang, Yuan Yao 0001, Hanghang Tong, Xuan Huo, Ming Li 0005, Feng Xu 0007, Jian Lu 0001 |
Knowl. Inf. Syst. | 5 |
| 2019 | Automatic Code Review by Learning the Revision of Source CodeabstractCode review is the process of manual inspection on the revision of the source code in order to find out whether the revised source code eventually meets the revision requirements. However, manual code review is time-consuming, and automating such the code review process will alleviate the burden of code reviewers and speed up the software maintenance process. To construct the model for automatic code review, the characteristics of the revisions of source code (i.e., the difference between the two pieces of source code) should be properly captured and modeled. Unfortunately, most of the existing techniques can easily model the overall correlation between two pieces of source code, but not for the “difference” between two pieces of source code. In this paper, we propose a novel deep model named DACE for automatic code review. Such a model is able to learn revision features by contrasting the revised hunks from the original and revised source code with respect to the code context containing the hunks. Experimental results on six open source software projects indicate by learning the revision features, DACE can outperform the competing approaches in automatic code review. Ming Li 0005, David Lo 0001, Ferdian Thung, Xuan Huo |
AAAI | 2 |
| 2019 | Find Me if You Can: Deep Software Clone Detection by Exploiting the Contest between the Plagiarist and the DetectorabstractCode clone is common in software development, which usually leads to software defects or copyright infringement. Researchers have paid significant attention to code clone detection, and many methods have been proposed. However, the patterns for generating the code clones do not always remain the same. In order to fool the clone detection systems, the plagiarists, known as the clone creator, usually conduct a series of tricky modifications on the code fragments to make the clone difficult to detect. The existing clone detection approaches, which neglects the dynamics of the “contest” between the plagiarist and the detectors, is doomed to be not robust to adversarial revision of the code. In this paper, we propose a novel clone detection approach, namely ACD, to mimic the adversarial process between the plagiarist and the detector, which enables us to not only build strong a clone detector but also model the behavior of the plagiarists. Such a plagiarist model may in turn help to understand the vulnerability of the current software clone detection tools. Experiments show that the learned policy of plagiarist can help us build stronger clone detector, which outperforms the existing clone detection methods. Yan-Ya Zhang, Ming Li 0005 |
AAAI | 2 |
| 2019 | Learning Uniform Semantic Features for Natural Language and Programming Language Globally, Locally and SequentiallyabstractSemantic feature learning for natural language and programming language is a preliminary step in addressing many software mining tasks. Many existing methods leverage information in lexicon and syntax to learn features for textual data. However, such information is inadequate to represent the entire semantics in either text sentence or code snippet. This motivates us to propose a new approach to learn semantic features for both languages, through extracting three levels of information, namely global, local and sequential information, from textual data. For tasks involving both modalities, we project the data of both types into a uniform feature space so that the complementary knowledge in between can be utilized in their representation. In this paper, we build a novel and general-purpose feature learning framework called UniEmbed, to uniformly learn comprehensive semantic representation for both natural language and programming language. Experimental results on three real-world software mining tasks show that UniEmbed outperforms state-of-the-art models in feature learning and prove the capacity and effectiveness of our model. Wenhao Zheng 0001, Ming Li 0005 |
AAAI | 3 |
| 2019 | On the Robust Splitting Criterion of Random ForestabstractSplitting criteria have played an important role in the construction of decision trees, and various trees have been developed based on different criteria. This work presents a unified framework on various splitting criteria from the perspective of loss functions, and most classical splitting criteria can be viewed essentially as the optimizations of loss functions in this framework. We further introduce a new splitting criterion, named pairwise gain, which is motivated from a lower bound on the mutual coupling of pairwise loss. Theoretically, we prove that this new criterion is robust to symmetric and asymmetric label noises simultaneously. Based on this new criterion, we develop another variant of random forests, and extensive experiments are provided to verify its robustness. Bin-Bin Yang, Wei Gao 0008, Ming Li 0005 |
ICDM | 3 |
| 2019 | Recurrent Aggregation Learning for Multi-view Echocardiographic Sequences Segmentation
Ming Li 0005, Weiwei Zhang 0006, Guang Yang 0006, Chengjia Wang, Heye Zhang, Huafeng Liu 0003, Shuo Li 0001 |
MICCAI (2) | 1 |
| 2019 | Towards One Reusable Model for Various Software Defect Mining Tasks
Heng-Yi Li, Ming Li 0005, Zhi-Hua Zhou |
PAKDD (3) | 2 |
| 2019 | DeepReview: Automatic Code Review Using Deep Multi-instance Learning
Heng-Yi Li, Ferdian Thung, Xuan Huo, Ming Li 0005, David Lo 0001 |
PAKDD (2) | 6 |
| 2019 | CodeAttention: translating source code to comments by exploiting the code constructs
Wenhao Zheng 0001, Ming Li 0005, Jianxin Wu 0001 |
Frontiers Comput. Sci. | 3 |
| 2019 | On cost-effective software defect prediction: Classification or ranking?
Xuan Huo, Ming Li 0005 |
Neurocomputing | 2 |
| 2019 | Distributed Deep Forest and its Application to Automatic Detection of Cash-Out FraudabstractInternet companies are facing the need for handling large-scale machine learning applications on a daily basis and distributed implementation of machine learning algorithms which can handle extra-large-scale tasks with great performance is widely needed. Deep forest is a recently proposed deep learning framework which uses tree ensembles as its building blocks and it has achieved highly competitive results on various domains of tasks. However, it has not been tested on extremely large-scale tasks. In this work, based on our parameter server system, we developed the distributed version of deep forest. To meet the need for real-world tasks, many improvements are introduced to the original deep forest model, including MART (Multiple Additive Regression Tree) as base learners for efficiency and effectiveness consideration, the cost-based method for handling prevalent class-imbalanced data, MART based feature selection for high dimension data, and different evaluation metrics for automatically determining the cascade level. We tested the deep forest model on an extra-large-scale task, i.e., automatic detection of cash-out fraud, with more than 100 million training samples. Experimental results showed that the deep forest model has the best performance according to the evaluation metrics from different perspectives even with very little effort for parameter tuning. This model can block fraud transactions in a large amount of money each day. Even compared with the best-deployed model, the deep forest model can additionally bring a significant decrease in economic loss each day. Ya-Lin Zhang 0001, Jun Zhou 0011, Wenhao Zheng 0001, Ji Feng, Ming Li 0005, Zhiqiang Zhang 0012, Chaochao Chen 0001, Xiaolong Li 0005, Yuan Qi 0001, Zhi-Hua Zhou |
ACM Trans. Intell. Syst. Technol. | 7 |
| 2018 | Semi-Supervised AUC Optimization Without Guessing Labels of Unlabeled DataabstractSemi-supervised learning, which aims to construct learners that automatically exploit the large amount of unlabeled data in addition to the limited labeled data, has been widely applied in many real-world applications. AUC is a well-known performance measure for a learner, and directly optimizing AUC may result in a better prediction performance. Thus, semi-supervised AUC optimization has drawn much attention. Existing semi-supervised AUC optimization methods exploit unlabeled data by explicitly or implicitly estimating the possible labels of the unlabeled data based on various distributional assumptions. However, these assumptions may be violated in many real-world applications, and estimating labels based on the violated assumption may lead to poor performance. In this paper, we argue that, in semi-supervised AUC optimization, it is unnecessary to guess the possible labels of the unlabeled data or prior probability based on any distributional assumptions. We analytically show that the AUC risk can be estimated unbiasedly by simply treating the unlabeled data as both positive and negative. Based on this finding, two semi-supervised AUC optimization methods named Samult and Sampura are proposed. Experimental results indicate that the proposed methods outperform the existing methods. Zheng Xie 0001, Ming Li 0005 |
AAAI | 2 |
| 2018 | Holistic and Deep Feature Pyramids for Saliency Detection
Shizhong Dong, Zhifan Gao, Shanhui Sun, Xin Wang 0045, Ming Li 0005, Heye Zhang, Guang Yang 0006, Huafeng Liu 0003, Shuo Li 0001 |
BMVC | 5 |
| 2018 | Learning Semantic Features for Software Defect Prediction by Code Comments EmbeddingabstractSoftware Quality Assurance (SQA) is essential in software development and many defect prediction methods based on machine learning have been proposed to identify defective modules. However, most existing defect prediction models do not provide good defect prediction results, and the semantic features reflecting the detective patterns may not be well-captured via traditional feature extraction methods. More information such as code comments should be also be embedded to generate semantic features respecting the source code functionality. Therefore, how to embed code comments for defect prediction is a big challenge, and another problem is that many comments of source code are missing in real-world applications. In this paper, we propose a novel defect prediction model named CAP-CNN (Convolutional Neural Network for Comments Augmented Programs), which is a deep learning model that automatically embeds code comments in generating semantic features from the source code for software defect prediction. To overcome the missing comments problem, a novel training strategy is used in CAP-CNN that the network encodes and absorb comments information to generate semantic features automatically during training process, which does not need testing modules to contain comments. Experimental results on several widely-used software data sets indicate that the comment features are able to improve defect prediction performance. Xuan Huo, Yang Yang 0074, Ming Li 0005, De-Chuan Zhan |
ICDM | 3 |
| 2018 | T2S: Domain Adaptation Via Model-Independent Inverse Mapping and Model ReuseabstractDomain adaptation, which is able to leverage the abundant supervision from the source domain and limited supervision in the target domain to construct a model for the data in the target domain, has drawn significant attentions. Most of the existing domain adaptation methods elaborate to map the information derived from the source domain to the target domain for model construction in the target domain. However, such a 'Source' (S) to 'Target' (T) mapping usually involves 'tailoring' the information from the source domain to fit the target domain, which may lose valuable information in the source domain for model construction. Moreover, such a mapping is usually tightly coupled with the model construction, which is more complex than a separate model construction or mapping construction. In this paper, we provide an alternative way for domain adaptation, named T2S. Instead of mapping the 'S' to 'T' and constructing a model in 'T', we inversely map 'T' to 'S' and reuse the model that has been well-trained with abundant information in 'S' for prediction. Such an approach enjoys the abundant information in source domain for model construction and the simplicity of learning mapping separately with limited supervision in target domain. Experiments on both synthetic and real-world data sets indicate the effectiveness of our framework. Zhi-Yu Shen, Ming Li 0005 |
ICDM | 2 |
| 2018 | Positive and Unlabeled Learning for Detecting Software Functional Clones with Adversarial TrainingabstractSoftware clone detection is an important problem for software maintenance and evolution and it has attracted lots of attentions. However, existing approaches ignore a fact that people would label the pairs of code fragments as \emph{clone} only if they happen to discover the clones while a huge number of undiscovered clone pairs and non-clone pairs are left unlabeled. In this paper, we argue that the clone detection task in the real-world should be formalized as a Positive-Unlabeled (PU) learning problem, and address this problem by proposing a novel positive and unlabeled learning approach, namely CDPU, to effectively detect software functional clones, i.e., pieces of codes with similar functionality but differing in both syntactical and lexical level, where adversarial training is employed to improve the robustness of the learned model to those non-clone pairs that look extremely similar but behave differently. Experiments on software clone detection benchmarks indicate that the proposed approach together with adversarial training outperforms the state-of-the-art approaches for software functional clone detection. Huihui Wei, Ming Li 0005 |
IJCAI | 2 |
| 2018 | Cutting the Software Building Efforts in Continuous Integration by Semi-Supervised Online AUC OptimizationabstractContinuous Integration (CI) systems aim to provide quick feedback on the success of the code changes by keeping on building the entire systems upon code changes are committed. However, building the entire software system is usually resource and time consuming. Thus, build outcome prediction is usually employed to distinguish the successful builds from the failed ones to cut the building efforts on those successful builds that do not result in any immediate action of the developer. Nevertheless, build outcome prediction in CI is challenging since the learner should be able to learn from a stream of build events with and without the build outcome labels and provide immediate prediction on the next build event. Also, the distribution of the successful and the failed builds are often highly imbalanced. Unfortunately, the existing methods fail to address these challenges well. In this paper, we address these challenges by proposing a semi-supervised online AUC optimization method for CI build outcome prediction. Experiments indicate that our method is able to cut the software building efforts by effectively identify the successful builds, and it outperforms the existing methods that elaborate to address part of these challenges. Zheng Xie 0001, Ming Li 0005 |
IJCAI | 2 |
| 2017 | Enhancing the Unified Features to Locate Buggy Files by Exploiting the Sequential Nature of Source CodeabstractBug reports provide an effective way for end-users to disclose potential bugs hidden in a software system, while automatically locating the potential buggy source files according to a bug report remains a great challenge in software maintenance. Many previous approaches represent bug reports and source code from lexical and structural information correlated their relevance by measuring their similarity, and recently a CNN-based model is proposed to learn the unified features for bug localization, which overcomes the difficulty in modeling natural and programming languages with different structural semantics. However, previous studies fail to capture the sequential nature of source code, which carries additional semantics beyond the lexical and structural terms and such information is vital in modeling program functionalities and behaviors. In this paper, we propose a novel model LS-CNN, which enhances the unified features by exploiting the sequential nature of source code. LS-CNN combines CNN and LSTM to extract semantic features for automatically identifying potential buggy source code according to a bug report. Experimental results on widely-used software projects indicate that LS-CNN significantly outperforms the state-of-the-art methods in locating buggy files. Xuan Huo, Ming Li 0005 |
IJCAI | 2 |
| 2017 | Supervised Deep Features for Software Functional Clone Detection by Exploiting Lexical and Syntactical Information in Source CodeabstractSoftware clone detection, aiming at identifying out code fragments with similar functionalities, has played an important role in software maintenance and evolution. Many clone detection approaches have been proposed. However, most of them represent source codes with hand-crafted features using lexical or syntactical information, or unsupervised deep features, which makes it difficult to detect the functional clone pairs, i.e., pieces of codes with similar functionality but differing in both syntactical and lexical level. In this paper, we address the software functional clone detection problem by learning supervised deep features. We formulate the clone detection as a supervised learning to hash problem and propose an end-to-end deep feature learning framework called CDLH for functional clone detection. Such framework learns hash codes by exploiting the lexical and syntactical information for fast computation of functional similarity between code fragments. Experiments on software clone detection benchmarks indicate that the CDLH approach is effective and outperforms the state-of-the-art approaches in software functional clone detection. Huihui Wei, Ming Li 0005 |
IJCAI | 2 |
| 2017 | Cost-effective build outcome prediction using cascaded classifiersabstractSoftware developers use continuous integration to find defects in the early stage and reduce risk. But this process can be resource and time consuming, which decreases the efficiency of development. In this work, we adopt cascaded classifiers to predict the build outcome and study what kinds of attributes are potentially useful for this process. We emphasize on the "failed" instances which bring more cost. Our experiments reveal that our approach outperforms other commonly used classifiers. It reduces 51.7% of the waiting time and server workload while identifying 85.2% of the defective builds. Ansong Ni, Ming Li 0005 |
MSR | 2 |
| 2017 | The best answer prediction by exploiting heterogeneous data on software development Q&A forum
Wenhao Zheng 0001, Ming Li 0005 |
Neurocomputing | 2 |
| 2016 | Learning Unified Features from Natural and Programming Languages for Locating Buggy Source Code
Xuan Huo, Ming Li 0005, Zhi-Hua Zhou |
IJCAI | 2 |
| 2015 | Constrained feature selection for localizing faultsabstractDevelopers often take much time and effort to find buggy program elements. To help developers debug, many past studies have proposed spectrum-based fault localization techniques. These techniques compare and contrast correct and faulty execution traces and highlight suspicious program elements. In this work, we propose constrained feature selection algorithms that we use to localize faults. Feature selection algorithms are commonly used to identify important features that are helpful for a classification task. By mapping an execution trace to a classification instance and a program element to a feature, we can transform fault localization to the feature selection problem. Unfortunately, existing feature selection algorithms do not perform too well, and we extend its performance by adding a constraint to the feature selection formulation based on a specific characteristic of the fault localization problem. We have performed experiments on a popular benchmark containing 154 faulty versions from 8 programs and demonstrate that several variants of our approach can outperform many fault localization techniques proposed in the literature. Using Wilcoxon rank-sum test and Cliff's d effect size, we also show that the improvements are both statistically significant and substantial. Tien-Duy B. Le, David Lo 0001, Ming Li 0005 |
ICSME | 3 |
| 2012 | Sample-based software defect prediction with active and semi-supervised learning
Ming Li 0005, Hongyu Zhang 0002, Rongxin Wu, Zhi-Hua Zhou |
Autom. Softw. Eng. | 1 |
| 2011 | Software Defect Detection with Rocus
Yuan Jiang 0001, Ming Li 0005, Zhi-Hua Zhou |
J. Comput. Sci. Technol. | 2 |
| 2010 | Exploiting remote learners in Internet environment with agents
Ming Li 0005, Wei Wang 0028, Zhi-Hua Zhou |
Sci. China Inf. Sci. | 1 |
| 2010 | Semi-supervised learning by disagreement
Zhi-Hua Zhou, Ming Li 0005 |
Knowl. Inf. Syst. | 2 |
| 2009 | Learning instance specific distances using metric propagationabstractIn many real-world applications, such as image retrieval, it would be natural to measure the distances from one instance to others using instance specific distance which captures the distinctions from the perspective of the concerned instance. However, there is no complete framework for learning instance specific distances since existing methods are incapable of learning such distances for test instance and unlabeled data. In this paper, we propose the Isd method to address this issue. The key of Isd is metric propagation, that is, propagating and adapting metrics of individual labeled examples to individual unlabeled instances. We formulate the problem into a convex optimization framework and derive efficient solutions. Experiments show that Isd can effectively learn instance specific distances for labeled as well as unlabeled instances. The metric propagation scheme can also be used in other scenarios. De-Chuan Zhan, Ming Li 0005, Yufeng Li 0008, Zhi-Hua Zhou |
ICML | 2 |
| 2009 | Exploiting Multi-Modal Interactions: A Unified Framework
Ming Li 0005, Xiao-Bing Xue, Zhi-Hua Zhou |
IJCAI | 1 |
| 2009 | Semi-supervised document retrieval
Ming Li 0005, Hang Li 0001, Zhi-Hua Zhou |
Inf. Process. Manag. | 1 |
| 2009 | Mining extremely small data sets with application to software reuseabstractAbstract A serious problem encountered by machine learning and data mining techniques in software engineering is the lack of sufficient data. For example, there are only 24 examples in the current largest data set on software reuse. In this paper, a recently proposed machine learning algorithm is modified for mining extremely small data sets. This algorithm works in a twice‐learning style. In detail, a random forest is trained from the original data set at first. Then, virtual examples are generated from the random forest and used to train a single decision tree. In contrast to the numerous discrepancies between the empirical data and expert opinions reported by previous research, our mining practice shows that the empirical data are actually consistent with expert opinions. Copyright © 2008 John Wiley & Sons, Ltd. Yuan Jiang 0001, Ming Li 0005, Zhi-Hua Zhou |
Softw. Pract. Exp. | 2 |
| 2008 | Mining Bulletin Board Systems Using Community Generation
Ming Li 0005, Zhongfei Zhang, Zhi-Hua Zhou |
PAKDD | 1 |
| 2008 | Online Manifold Regularization: A New Learning Setting and Empirical Study
Andrew B. Goldberg, Ming Li 0005, Xiaojin Zhu 0001 |
ECML/PKDD (1) | 2 |
| 2007 | Semisupervised Regression with Cotraining-Style AlgorithmsabstractThe traditional setting of supervised learning requires a large amount of labeled training examples in order to achieve good generalization. However, in many practical applications, unlabeled training examples are readily available, but labeled ones are fairly expensive to obtain. Therefore, semisupervised learning has attracted much attention. Previous research on semisupervised learning mainly focuses on semisupervised classification. Although regression is almost as important as classification, semisupervised regression is largely understudied. In particular, although cotraining is a main paradigm in semisupervised learning, few works has been devoted to cotraining-style semisupervised regression algorithms. In this paper, a cotraining-style semisupervised regression algorithm, that is, COREG, is proposed. This algorithm uses two regressors, each labels the unlabeled data for the other regressor, where the confidence in labeling an unlabeled example is estimated through the amount of reduction in mean squared error over the labeled neighborhood of that example. Analysis and experiments show that COREG can effectively exploit unlabeled data to improve regression estimates. Zhi-Hua Zhou, Ming Li 0005 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2007 | Improve Computer-Aided Diagnosis With Machine Learning Techniques Using Undiagnosed SamplesabstractIn computer-aided diagnosis (CAD), machine learning techniques have been widely applied to learn a hypothesis from diagnosed samples to assist the medical experts in making a diagnosis. To learn a well-performed hypothesis, a large amount of diagnosed samples are required. Although the samples can be easily collected from routine medical examinations, it is usually impossible for medical experts to make a diagnosis for each of the collected samples. If a hypothesis could be learned in the presence of a large amount of undiagnosed samples, the heavy burden on the medical experts could be released. In this paper, a new semisupervised learning algorithm named Co-Forest is proposed. It extends the co-training paradigm by using a well-known ensemble method named Random Forest, which enables Co-Forest to estimate the labeling confidence of undiagnosed samples and easily produce the final hypothesis. Experiments on benchmark data sets verify the effectiveness of the proposed algorithm. Case studies on three medical data sets and a successful application to microcalcification detection for breast cancer diagnosis show that undiagnosed samples are helpful in building CAD systems, and Co-Forest is able to enhance the performance of the hypothesis that is learned on only a small amount of diagnosed samples by utilizing the available undiagnosed samples. Ming Li 0005, Zhi-Hua Zhou |
IEEE Trans. Syst. Man Cybern. Part A | 1 |
| 2005 | Semi-Supervised Regression with Co-Training
Zhi-Hua Zhou, Ming Li 0005 |
IJCAI | 2 |
| 2005 | SETRED: Self-training with Editing
Ming Li 0005, Zhi-Hua Zhou |
PAKDD | 1 |
| 2005 | Multi-Instance Learning Based Web Mining
Zhi-Hua Zhou, Ming Li 0005 |
Appl. Intell. | 3 |
| 2005 | Tri-Training: Exploiting Unlabeled Data Using Three ClassifiersabstractIn many practical data mining applications, such as Web page classification, unlabeled training examples are readily available, but labeled ones are fairly expensive to obtain. Therefore, semi-supervised learning algorithms such as co-training have attracted much attention. In this paper, a new co-training style semi-supervised learning algorithm, named tri-training, is proposed. This algorithm generates three classifiers from the original labeled example set. These classifiers are then refined using unlabeled examples in the tri-training process. In detail, in each round of tri-training, an unlabeled example is labeled for a classifier if the other two classifiers agree on the labeling, under certain conditions. Since tri-training neither requires the instance space to be described with sufficient and redundant views nor does it put any constraints on the supervised learning algorithm, its applicability is broader than that of previous co-training style algorithms. Experiments on UCI data sets and application to the Web page classification task indicate that tri-training can effectively exploit unlabeled data to enhance the learning performance. Zhi-Hua Zhou, Ming Li 0005 |
IEEE Trans. Knowl. Data Eng. | 2 |