EDBT 2026 Demo / reviewers in the wild / expert
Hui Li 0014
dblp:66/3387-14
· DBLP profile ↗
51ranked-venue papers
5as first author
44since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 23 · 5 first-author · 19 since 2021Artificial intelligence and machine learning · 13 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ContrastKV: Robust KV Cache Eviction via Contrastive Signal Fusion for Multi-Query GeneralizationabstractLarge Language Models (LLMs) face significant memory and latency overheads during long-context inference due to the growing KV cache, especially in Knowledge Base Question Answering (KBQA) settings that require support for multiple downstream queries. Query-aware eviction methods do not generalize across queries, while existing query-agnostic approaches rely on a single proxy query, leading to fragile eviction decisions under high eviction ratios. We propose ContrastKV, a robust query-agnostic KV cache eviction algorithm for multi-query generalization. ContrastKV introduces a contrastive signal fusion mechanism that jointly exploits complementary semantic and non-semantic signals. By contrasting semantic consistency with structural robustness, the method constructs a more reliable eviction criterion that alleviates the blind spots of single-query proxies. The framework integrates efficient signal generation, parallel importance scoring, and multi-level fusion across heads and layers. Experiments show that ContrastKV outperforms state-of-the-art methods, retaining up to 92% accuracy with only 20% of the KV cache budget, while reducing decoding latency by approximately 50% and significantly lowering GPU memory usage. Xingchi Chen, Peiyuan Zong, Ziqiang Gao, Qing Li 0006, Yong Jiang 0001, Fa Zhu, Hui Li 0014 |
ACL (1) | 7 |
| 2026 | Modeling Relational Logic Circuits for and-Inverter Graph Convolutional NetworkabstractThe automation of logic circuit design enhances chip performance, energy efficiency, and reliability, and is widely applied in the field of Electronic Design Automation (EDA). And-Inverter Graphs (AIGs) efficiently represent, optimize, and verify the functional characteristics of digital circuits, enhancing the efficiency of EDA development. Due to the complex structure and large scale of nodes in real-world AIGs, accurate modeling is challenging, leading to existing work lacking the ability to jointly model functional and structural characteristics, as well as insufficient dynamic information propagation capability. To address the aforementioned challenges, we propose AIGer, with the aim to enhance the expression of AIGs and thereby improve the efficiency of EDA development. Specifically, AIGer consists of two components: 1) Node logic feature initialization embedding component and 2) AIGs feature learning network component. The node logic feature initialization embedding component projects logic nodes, such as AND and NOT, into independent semantic spaces, to enable effective node embedding for subsequent processing. Building upon this, the AIGs feature learning network component employs a heterogeneous graph convolutional network, designing dynamic relationship weight matrices and differentiated information aggregation approaches to better represent the original structure and information of AIGs. The combination of these two components enhances AIGer’s ability to jointly model functional and structural characteristics and improves its message passing capability, thereby strengthening its expressive power for AIGs and enhancing the development efficiency of logic circuits. Experimental results indicate that AIGer outperforms the current best models in the Signal Probability Prediction (SPP) task, improving MAE and MSE by 18.95% and 44.44%, respectively. In the Truth Table Distance Prediction (TTDP) task, AIGer achieves improvements of 33.57% and 14.79% in MAE and MSE, respectively, compared to the best-performing models1. Weihao Sun, Shikai Guo, Qian Ma 0003, Hui Li 0014, Yongpeng Weng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2026 | Cross Attention and Intra-Layer Attention in Heterogeneous Graph Neural Networks for Drug-Target Interaction PredictionabstractIn recent years, computational prediction of drug-target interactions (DTIs) has become essential for drug discovery and repositioning. However, traditional experimental approaches for DTI identification are time-consuming and costly. To address this, many machine learning-based methods have been developed, yet most existing models neglect important information interaction between drugs and targets in drug-target pairs (DTPs) during drug-target interaction. In this study, we propose a novel cross-attention and intra-layer attention mechanism within a heterogeneous graph neural network (CAIHGNN) for DTI prediction. The cross-attention mechanism allows for dynamic learning of feature correlations between drugs and targets, while the intra-layer attention captures both explicit and implicit interactions within DTPs. Additionally, we introduce a drug-target pair correlation graph to exploit high-order interactions between DTPs. Extensive experiments on two biological heterogeneous datasets demonstrate the superior performance of our proposed method in accurately predicting DTIs. Furthermore, the model exhibits robust generalization in case study, showing promise for real-world drug discovery applications. Kuiyang Che, Xirun Wei, Hui Li 0014, Shikai Guo |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2026 | Molecular-Driven Multi-View Hypergraph Contrastive Learning for Drug-Drug Interaction PredictionabstractRecent concerns have arisen over adverse reactions caused by drug combinations, and drug-drug interaction (DDI) prediction helps identify potential risks by forecasting interactions between drugs. Previous methods have primarily explored drug interactions from the superficial level of drug molecules, often overlooking the internal structural information of the molecules. To this end, we propose Mol-HCL, a multi-view hypergraph contrastive learning framework based on molecular view. In this framework, we construct the molecular view to learn the internal information of drug molecules and, based on this, develop structural view and semantic view. These three views collaboratively learn both intra-molecular and inter-molecular information. Subsequently, we incorporate hypernodes into the structural view and design a novel hyperchain, integrating it into the semantic view to capture latent neighbor drug node structural relationships and long-range DDI chain semantic information. After that, contrastive learning is performed between the structural hypergraph and the molecular view, as well as between the semantic hypergraph and the molecular view, to enhance the representations learned from the molecular view. Finally, we conduct experiments on two real-world scientific datasets. The experimental results demonstrate a significant improvement of Mol-HCL over existing methods, showcasing its effectiveness and advantages in DDI prediction. Shikai Guo, Hui Li 0014, Qian Ma 0003 |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2026 | SMem-Diff: A Simple Memory-Augmented Diffusion Model for Effective Video Deblurring on Cloud-Edge Servers
Qichuang Liu, Hui Li 0014, Fa Zhu, Xingchi Chen, Qing Li 0006, Moustafa Youssef 0001, Giancarlo Fortino |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Estimating Uncertainty in Line-Level Defect Prediction via Perceptual Borderline OversamplingabstractSoftware defect prediction aims to identify potentially defective software modules using various techniques, while fine-grained line-level defect prediction can pinpoint defective lines of code. This helps developers promptly discover and fix errors, thereby enhancing the efficiency of testing and code review. However, previous studies often overlook the impact of characterizing noise and the skewed distribution of defect knowledge in software projects, making it difficult for current methods to achieve satisfactory accuracy and cost-effectiveness in software defect prediction. To address these challenges, we propose a model named EU-LLDP, which effectively resolves the issue of low cost-effectiveness in line-level defect prediction models. Specifically, the EU-LLDP model consists of two main components: the defect mining component mines the most valuable defect knowledge from numerous software defects using prediction probability matrices, noise labels, and the borderline information of code vectorizations. The adaptive resampling component samples valuable defect knowledge through the density distribution of defect knowledge, thereby making full use of existing defect knowledge and improving the cost-effectiveness of line-level software defect prediction models. Seven comprehensive experiments were conducted on 32 defect datasets from 9 Java open source systems using file-level prediction models and line-level defect prediction models to evaluate the effectiveness of the EU-LLDP model. The EU-LLDP model improves the state-of-the-art file-level defect prediction model in terms of Balanced Accuracy by 9.87%, the MCC by 38.09%, and enhances the state-of-the-art line-level defect prediction method in terms of Recall@Top20%LOC by 44.16%, and Effort@Top20%Recall by 17.62%. These results fully demonstrate the effectiveness of EU-LLDP in improving the accuracy and cost-effectiveness of Software defect prediction. Shikai Guo, Hui Li 0014, Rong Chen 0003 |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2025 | LBMKGC: Large Model-Driven Balanced Multimodal Knowledge Graph CompletionabstractMulti-modal Knowledge Graph Completion (MMKGC) aims to predict missing entities, relations, or attributes in knowledge graphs by collaboratively modeling the triple structure and multimodal information (e.g., text, images, videos) associated with entities.
This approach facilitates the automatic discovery of previously unobserved factual knowledge.
However, existing MMKGC methods encounter several critical challenges: (i) the imbalance of inter-entity information across different modalities; (ii) the heterogeneity of intra-entity multimodal information; and (iii) for a given entity, the informational contributions of different modalities are inconsistent across contexts.
In this paper, we propose a novel **L**arge model-driven **B**alanced **M**ultimodal **K**nowledge **G**raph **C**ompletion framework, termed LBMKGC.
Subsequently, to bridge the semantic gap between heterogeneous modalities, LBMKGC aligns the multimodal embeddings of entities semantically by using the CLIP (Contrastive Language-Image Pre-Training) model.
Furthermore, LBMKGC adaptively fuses multimodal embeddings with relational guidance by distinguishing between the perceptual and conceptual attributes of triples.
Finally, extensive experiments conducted against 21 state-of-the-art baselines demonstrate that LBMKGC achieves superior performance across diverse datasets and scenarios while maintaining efficiency and generalizability.
Our code and data are publicly available at: https://github.com/guoynow/LBMKGC. Qian Ma 0003, Hui Li 0014, Furui Zhan, Yu Gu 0002, Ge Yu 0001, Shikai Guo |
NeurIPS | 3 |
| 2025 | Sul-BertGRU: an ensemble deep learning method integrating information entropy-enhanced BERT and directional multi-GRU for S-sulfhydration sites predictionabstractMOTIVATION: S-sulfhydration, a crucial post-translational protein modification, is pivotal in cellular recognition, signaling processes, and the development and progression of cardiovascular and neurological disorders, so identifying S-sulfhydration sites is crucial for studies in cell biology. Deep learning shows high efficiency and accuracy in identifying protein sites compared to traditional methods that often lack sensitivity and specificity in accurately locating nonsulfhydration sites. Therefore, we employ deep learning methods to tackle the challenge of pinpointing S-sulfhydration sites. RESULTS: In this work, we introduce a deep learning approach called Sul-BertGRU, designed specifically for predicting S-sulfhydration sites in proteins, which integrates multi-directional gated recurrent unit (GRU) and BERT. First, Sul-BertGRU proposes an information entropy-enhanced BERT (IE-BERT) to preprocess protein sequences and extract initial features. Subsequently, confidence learning is employed to eliminate potential S-sulfhydration samples from the nonsulfhydration samples and select reliable negative samples. Then, considering the directional nature of the modification process, protein sequences are categorized into left, right, and full sequences centered on cysteines. We build a multi-directional GRU to enhance the extraction of directional sequence features and model the details of the enzymatic reaction involved in S-sulfhydration. Ultimately, we apply a parallel multi-head self-attention mechanism alongside a convolutional neural network to deeply analyze sequence features that might be missed at a local level. Sul-BertGRU achieves sensitivity, specificity, precision, accuracy, Matthews correlation coefficient, and area under the curve scores of 85.82%, 68.24%, 74.80%, 77.44%, 55.13%, and 77.03%, respectively. Sul-BertGRU demonstrates exceptional performance and proves to be a reliable method for predicting protein S-sulfhydration sites. AVAILABILITY AND IMPLEMENTATION: The source code and data are available at https://github.com/Severus0902/Sul-BertGRU/. Xirun Wei, Kuiyang Che, Hui Li 0014, Shikai Guo |
Bioinform. | 5 |
| 2025 | Structuring Semantic-Aware Relations Between Bugs and Patches for Accurate Patch EvaluationabstractABSTRACT Patches can help fix security vulnerabilities and optimize software performance, thereby enhancing the quality and security of the software. Unfortunately, patches generated by automated program repair tools are not always correct, as they may introduce new bugs or fail to fully rectify the original issue. Various methods for evaluating patch correctness have been proposed. However, most methods face the challenge of capturing long‐distance dependencies in patch correctness evaluation, which leads to a decline in the predictive performance of the models. To address the challenge, this paper presents a method named Qamhaen to evaluate the correctness of patches generated by APR. Specifically, text embedding of bugs and patches component address the challenge of long‐distance dependencies across functions in patch correctness evaluation by using bug reports and patch descriptions as inputs instead of code snippets. BERT is employed for pretraining to capture these dependencies, followed by an additional multihead self‐attention mechanism for further feature extraction. Similarity evaluator component devises a similarity calculation to assess the effectiveness of patch descriptions in resolving issues outlined in bug reports. Comprehensive experiments are conducted on a dataset containing 9135 patches and a patch correctness assessment metric, and extensive experiments demonstrate that Qamhaen outperforms baseline methods in terms of overall performance across AUC, F1, +Recall, ‐Recall, and Precision. For example, compared to the baseline, Qamhaen achieves an F1 of 0.691, representing improvements of 24.2%, 22.1%, and 6.3% over the baseline methods, respectively. Hui Li 0014, Yongqian Chen, Xiaowei Pan, Shikai Guo |
J. Softw. Evol. Process. | 2 |
| 2025 | Dual-Network Cross-Learning for Metabolite-Disease Association PredictionabstractIn recent years, increasing evidence has demonstrated a close association between metabolites and various complex human diseases, providing valuable insights for disease diagnosis, treatment, and prevention. Although deep learning-based approaches have achieved certain success in predicting metabolic disease associations, challenges remain in enriching graph information and effectively integrating metabolic and disease features. To address these issues, this paper proposes a model named DCMDA, which extracts deep features of both metabolites and diseases using Dual-network Cross-learning for Metabolite-Disease Association prediction. DCMDA consists of three parts. The data processing module integrates similarity networks with association networks to construct a heterogeneous network. The feature extraction module extracts features from the metabolite-disease association network based on the non-negative matrix factorization method and from the heterogeneous network using graph autoencoder techniques. The feature fusion module combines the association matrix feature with the heterogeneous network feature through a Cross-Attention mechanism, thereby obtaining deep representations of metabolites and diseases. These features are then used to train the model to predict association scores between metabolites and diseases. Experimental results demonstrate that in 5-fold cross-validation, DCMDA achieves an area under the receiver operating characteristic curve (AUC) of 97.8% and an area under the precision-recall curve (AUPR) of 97.9%, outperforming state-of-the-art prediction methods. Yanxin Chen, Hui Li 0014, Shikai Guo |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2025 | MHMDA: "Similarity-Association-Similarity" Metapaths and Heterogeneous-Hyper Network Learning for MiRNA-Disease Association PredictionabstractIn recent years, microRNA (miRNA) has been recognized as crucial in the progression of human diseases. However, existing computational methods for identifying miRNA-disease associations often overlook the rich association information contained in specific long-distance pathways and lack effective exploration of potential associations. In this study, we propose a biologically interpretable "similarity-association-similarity" metapath and heterogeneous-hyper network (HeteroHyperNet) learning approach for miRNA-disease association prediction (MHMDA). In MHMDA, a "similarity-association-similarity" multi-hop metapaths learning method based on hierarchical attention perception is proposed to explore specific long-distance associated pathway information connecting potentially associated miRNAs and diseases. In addition, a HeteroHyperNet learning approach integrating heterogeneous network and hyper network is designed to progressively learn direct association information and potential association information between miRNA and disease. The "similarity-association-similarity" metapath with hierarchical attention significantly enhances the learning of long-distance biological associations, while the HeteroHyperNet comprehensively learns the known and potential associations of miRNA-disease, greatly improving the richness and accuracy of information. A large number of experimental results show that MHMDA has demonstrated excellent performance in the prediction of miRNA-disease association. In addition, cross independent dataset experiment and cold start experiment on miRNA and disease prove the effectiveness of MHMDA on sparse association points, and its stability and reliability in predicting potential miRNA-disease association are further confirmed. Yaomiao Zhao, Shikai Guo, Hui Li 0014 |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2025 | Prediction of miRNA-Disease Association Based on Biodiversity Association NetworkabstractMiRNA-disease association identification is of great significance to the development of clinical medicine and drug research. Present computational methods didn't consider rich biological information, such as the expression level changes of disease-related miRNAs and the association information between miRNA, disease and other types of biological entities. In this study, we propose a new method for prediction of MiRNA-Disease Association based on Biodiversity Association Network (BANMDA). BANMDA first collects multiple types of association information from multiple sources, including diseases, miRNAs and lncRNAs associations, expression level changes of disease-related miRNAs, miRNAs sequence information, and disease semantic information. Second, BANMDA extracts diversity association features and diversity biological features based on two heterogeneous graph structure to represent miRNAs and diseases at multiple levels. In diversity association module, edges are classified according to the expression level changes of disease-related miRNAs and the similarities between miRNAs and diseases. In diversity node module, lncRNAs associated with miRNA-diseases are collected to construct heterogeneous network and we propose an improved GCN algorithm to directly aggregate the higher-order neighborhood information in the heterogeneous graph. Finally, the bilinear decoder is applied to predict associations between miRNAs and diseases. Experimental results show that BANMDA can be used as a powerful tool to identify miRNA-disease associations. Chaorui Guo, Hui Li 0014, Shikai Guo |
IEEE Trans. Comput. Biol. Bioinform. | 5 |
| 2025 | Making Fault Localization in Online Service Systems More Actionable and InterpretableabstractOnline service systems struggle with accurately and quickly pinpointing and resolving failures within their intricate systems, and it therefore emerges the solutions for fault localization in the code. However, the previous fault localization models suffer from low localization accuracy and poor interpretability due to the complex dependencies among fault characteristics in industrial practice. To address this issue, challenges brought by the long-distance dependencies among fault features and the unbalanced distribution of fault knowledge, and to improve the interpretability of the model, we present a fault localization model in online service systems more actionable and interpretable, named FL-AIer. Specifically, FL-AIer consists of two components: the feature encoding component and the fault localization component. The feature encoding component utilizes graph attention networks to capture the complex spatio-temporal dependencies within fault features. Then, the fault localization component adopts a three-stage approach, leveraging a multi-attention mechanism to identify and prioritize the most relevant fault features for precise localization. Additionally, the Fault Knowledge Balancing module it contains introduces a weighted Kullback-Leibler divergence loss function to ensure that the model pays adequate attention to all fault features, addressing the issue of imbalanced fault knowledge distribution and enhancing localization performance. We conducted extensive experiments on four datasets, and the results demonstrated that FL-AIer effectively addressing the challenges of fault localization in online system environments, and consistently outperforms the state-of-the-art methods across various evaluation metrics such as A@1, A@2, A@3, A@5, and MAR. For instance, FL-AIer achieves significant improvements of 5.82%, 10.77%, 4.20%, and 15.56% on the A@1 metric, respectively. These results fully demonstrate the excellent effectiveness of FL-AIer in effectively addressing the challenges of fault localization in online system environments, surpassing the performance of existing state-of-the-art methods. Ke Xv, Shikai Guo, Hui Li 0014, Rong Chen 0003, He Jiang 0001 |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2025 | Context-based Transfer Learning for Structuring Fault Localization and Program Repair AutomationabstractAutomated software debugging plays a crucial role in aiding software developers to swiftly identify and attempt to rectify faults, thereby significantly reducing developers’ workload. Previous researches have predominantly relied on simplistic semantic deep learning or statistical analysis methods to locate faulty statements in diverse projects. However, code repositories often consist of lengthy sequences with long-distance dependencies, posing challenges for accurately modeling fault localization using these methods. In addition, the lack of joint reasoning among various faults prevents existing models from deeply capturing fault information. To address these challenges, we propose a method named CodeHealer to achieve accurate fault localization and program repair. CodeHealer comprises three components: a Deep Semantic Information Extraction Component that effectively extracts deep semantic features from suspicious code statements using classifiers based on Joint-attention mechanisms; a Suspicious Statement Ranking Component that combines various fault localization features and employs multilayer perceptrons to derive multidimensional vectors of suspicion values; and a Fault Repair Component that, based on ranked suspicious statements generated by fault localization, adopts a top-down approach using multiple classifiers based on Co-teaching mechanisms to select repair templates and generate patches. The experimental results indicate that when applied to fault localization, CodeHealer outperforms the best baseline method with improvements of 11.4%, 2.7%, and 1.6% on Top-1/3/5 metrics, respectively. It also reduces the MFR and MAR by 9.8% and 2.1%, where lower values denote better fault localization effectiveness. Additionally, in automated software debugging, CodeHealer fixes an additional 6 faults compared to the current best method, totaling 53 faults repaired. Lehuan Zhang, Shikai Guo, Hui Li 0014, Yu Chai, Rong Chen 0003, He Jiang 0001 |
ACM Trans. Softw. Eng. Methodol. | 4 |
| 2025 | BR-Hunter: Detect Information Types of Bug Reports From Online Community DiscussionsabstractIn community-based software development, live-chatting services are increasingly used to discuss bugs encountered during development. Many methods have emerged to identify bugs and produce bug reports, which further improve the efficiency of software development. However, previous methods still face challenges in understanding complex conversational structures and classifying sentences in bug reports, as entertaining or meaningless utterances often lower the quality of constructed bug reports. To address this issue, we propose a method named BR-Hunter, which comprises the following four components. Specifically, the data preprocessing component disentangles and denoises the live chats, while the utterance embedding component aims to extract the semantic features of each utterance in the conversations. The bug report identification component then models the conversation as a feature graph and uses Graph Neural Networks to identify conversations containing bug reports, thereby solving Challenge 1. Finally, the bug report synthesis (BRS) component tackles Challenge 2 by classifying and reassembling sentences from conversations containing bug reports, leveraging fine-tuned BERT and prompt learning techniques. Extensive experiments conducted on eight open source projects demonstrate that BR-Hunter achieves high accuracy in identifying bug reports. Compared to baseline methods, BR-Hunter improves the average F1 score by 36.41%, 24.80%, 68.92%, 46.77%, 52.84%, 25.80%, 25.25%, and 4.19%, respectively. And BR-Hunter also achieves an average improvement of 10.34% on the BRS task, compared with the state-of-the-art method. Huijiang Liu, Junyu Xiong, Shikai Guo, Hui Li 0014 |
IEEE Trans. Reliab. | 4 |
| 2025 | Extracting Meaningful Issue-Solution Pair From Collaborative Developer Live ChatsabstractThe live chats of developers often contain meaningful information in the form of issue–solution pairs. The issue–solution pairs can offer helpful references to others who seek solutions for the similar issues, which can improve software development efficiency by facilitating issue solving. However, previous approaches such as ISPY still struggle with unsatisfactory extraction accuracy, due to the entanglement and complexity of issue-solution pairs' feature information. To address these challenges, we propose an approach namedIS-Hunterfor mining issue-solution pairs from real-time chat data. Specifically,IS-Hunterconsists of four main components: the data preprocessing component disentangles and denoises raw chat logs, the utterance embedding component embeds utterances into vectors that subsequent components can easily process, the feature extraction component obtains textual, heuristic, and contextual feature that determines whether an utterance is topic-relevant, and the issue–solution pair prediction component predicts the utterance whether is an issue or a solution. The experimental results show that the performance of IS-Hunter outperforms the baseline methods in issue-detection and solution-extraction in terms of Precision, Recall, and F1-score. Compared with baseline methods, in issue-detection, IS-Hunter, respectively, achieves an average precision, recall, and F1-score of 0.74, 0.74, and 0.74, and it marks an obvious 4.23% improvement over the state-of-the-art approaches. Simultaneously, in solution-extraction, IS-Hunter achieves an average precision, recall, and F1-score of 0.83, 0.90, and 0.86 which is 4.88% higher than the best baseline methods. Jiawen Shen, Shikai Guo, Longfeng Chen, Hui Li 0014 |
IEEE Trans. Reliab. | 5 |
| 2025 | Line-Level Defect Prediction by Capturing Code Contexts With Graph Convolutional NetworksabstractSoftware defect prediction refers to the systematic analysis and review of software using various approaches and tools to identify potential defects or errors. Software defect prediction aids developers in swiftly identifying defects and optimizing development resource allocation, thus enhancing software quality and reliability. Previous defect prediction approaches still face two main limitations: 1) lacking of contextual semantic information and 2) Ignoring the joint reasoning between different granularities of defect predictions. In response to these challenges, we propose LineDef, a line-level defect prediction approach by capturing code contexts with graph convolutional networks. Specifically, LineDef comprises three components: the token embedding component, the graph extraction component, and the multi-granularity defect prediction component. The token embedding component maps each token to a vector to obtain a high-dimensional semantic feature representation of the token. Subsequently, the graph extraction component utilizes a sliding window to extract line-level and token-level graphs, addressing the challenge of capturing contextual semantic relationships in the code. Finally, the multi-granularity defect prediction component leverages graph convolutional layers and attention mechanisms to acquire prediction labels and risk scores, thereby achieving file-level and line-level defect prediction. Experimental studies on 32 datasets across 9 different software projects show that LineDef exhibits significantly enhanced balanced accuracy, ranging from 15.61% to 45.20%, compared to state-of-the-art file-level defect prediction approaches, and a remarkable cost-effectiveness improvement ranging from 15.32% to 278%, compared to state-of-the-art line-level defect prediction approaches. These results demonstrate that LineDef approach can extract more comprehensive information from lines of code for defect prediction. Shouyu Yin, Shikai Guo, Hui Li 0014, Rong Chen 0003, He Jiang 0001 |
IEEE Trans. Software Eng. | 3 |
| 2025 | Point cloud upsampling via a coarse-to-fine network with transformer-encoder
Yixi Li, Yanzhe Liu, Rong Chen 0003, Hui Li 0014 |
Vis. Comput. | 4 |
| 2024 | VF-Detector: Making Multi-Granularity Code Changes on Vulnerability Fix Detector Robust to Mislabeled Changes
Zhenkan Fu, Shikai Guo, Hui Li 0014, Rong Chen 0003, He Jiang 0001 |
IJCAI | 3 |
| 2024 | Nukplex: An Efficient Local Search Algorithm for Maximum K-Plex Problem
Yiyuan Wang 0002, Shimao Wang, Hui Li 0014, Ximing Li 0002, Minghao Yin |
IJCAI | 4 |
| 2024 | Deep Just-In-Time Defect Prediction Based on Double-Source Input Self-Attention Mechanism (S)abstractEnsuring high-quality software products is an important issue for software productions, it emerged Just-In-Time Quality Assurance to automatically identify potentially defective code as early as possible in recent years.The presented models mainly utilize Convolutional Neural Networks for automatic defect feature extraction.However, these models ignore the utilization of contextual information and are not suitable for large-scale projects.To address these issues, we propose a model named Multi-head Convolution Structure of Attention-JIT (short for MCSA-JIT), which comprises four layers.Specifically, Commit Message-Code Change(short for CM-CC) Text Feature Extraction Layer leverages a variant of the self-attention mechanism to extract features from both commit messages and code changes, CM-CC Structure Feature Extraction Layer captures the content information of the submissions and different CNN structures are constructed to extract structural information, Feature Combination Layer combines the text feature and structure feature to predict defects and output layer outputs the final value.Experimental results conducted on two software projects, QT and OPENSTACK, demonstrate that the best variant of MCSA-JIT achieves a relative improvement of 7.46% in terms of AUC on the OPENSTACK project and a relative improvement of 8.85% on the QT project when compared to state-of-the-art methods with the best performance. Weixiang Hong 0002, Hui Li 0014, Shikai Guo |
SEKE | 3 |
| 2024 | Structuring Meaningful Code Review Automation in Developer Community
Zhenzhen Cao, Sijia Lv, Hui Li 0014, Qian Ma 0003, Cheng Guo 0001, Shikai Guo |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | A local search algorithm with movement gap and adaptive configuration checking for the maximum weighted s-plex problem
Shuli Hu, Yiyuan Wang 0002, Minghao Yin, Hui Li 0014 |
Eng. Appl. Artif. Intell. | 7 |
| 2024 | Graph Confident Learning for Software Vulnerability Detection
Qian Wang 0034, Zhengdao Li, Hetong Liang, Xiaowei Pan, Hui Li 0014, Shikai Guo |
Eng. Appl. Artif. Intell. | 5 |
| 2024 | Detect software vulnerabilities with weight biases via graph neural networks
Huijiang Liu, Shuirou Jiang, Xuexin Qi, Hui Li 0014, Cheng Guo 0001, Shikai Guo |
Expert Syst. Appl. | 5 |
| 2024 | Automated patch correctness predicting to fix software defect
Zelong Zheng, Zijian Tao, Hui Li 0014, Shikai Guo |
Expert Syst. Appl. | 4 |
| 2024 | Context-based transfer learning for low resource code summarizationabstractAbstract Source code summaries improve the readability and intelligibility of code, help developers understand programs, and improve the efficiency of software maintenance and upgrade processes. Unfortunately, these code comments are often mismatched, missing, or outdated in software projects, resulting in developers needing to infer functionality from source code, affecting the efficiency of software maintenance and evolution. Various methods based on neuronal networks are proposed to solve the problem of synthesis of source code. However, the current work is being carried out on resource‐rich programming languages such as Java and Python, and some low‐resource languages may not perform well. In order to solve the above challenges, we propose a context‐based transfer learning model for low resource code summarization (LRCS), which learns the common information from the language with rich resources, and then transfers it to the target language model for further learning. It consists of two components: the summary generation component is used to learn the syntactic and semantic information of the code, and the learning transfer component is used to improve the generalization ability of the model in the learning process of cross‐language code summarization. Experimental results show that LRCS outperforms baseline methods in code summarization in terms of sentence‐level BLEU, corpus‐level BLEU and METEOR. For example, LRCS improves corpus‐level BLEU scores by 52.90%, 41.10%, and 14.97%, respectively, compared to baseline methods. Yu Chai, Lehuan Zhang, Hui Li 0014, Mengzhi Luo, Shikai Guo |
Softw. Pract. Exp. | 4 |
| 2024 | Estimating Uncertainty in Labeled Changes by SZZ Tools on Just-In-Time Defect PredictionabstractThe aim of Just-In-Time (JIT) defect prediction is to predict software changes that are prone to defects in a project in a timely manner, thereby improving the efficiency of software development and ensuring software quality. Identifying changes that introduce bugs is a critical task in just-in-time defect prediction, and researchers have introduced the SZZ approach and its variants to label these changes. However, it has been shown that different SZZ algorithms introduce noise to the dataset to a certain extent, which may reduce the predictive performance of the model. To address this limitation, we propose the Confident Learning Imbalance (CLI) model. The model identifies and excludes samples whose labels may be corrupted by estimating the joint distribution of noisy labels and true labels, and mitigates the impact of noisy data on the performance of the prediction model. The CLI consists of two components: identifying noisy data (Confident Learning Component) and generating a predicted probability matrix for imbalanced data (Imbalanced Data Probabilistic Prediction Component). The IDPP component generates precise predicted probabilities for each instance in the training set, while the CL component uses the generated predicted probability matrix and noise labels to clean up the noise and build a classification model. We evaluate the performance of our model through extensive experiments on a total of 126,526 changes from ten Apache open source projects, and the results show that our model outperforms the baseline methods. Shikai Guo, Sijia Lv, Rong Chen 0003, Hui Li 0014, He Jiang 0001 |
ACM Trans. Softw. Eng. Methodol. | 6 |
| 2024 | Analyzing and Detecting Information Types of Developer Live Chat ThreadsabstractOnline chatrooms serve as vital platforms for information exchange among software developers. With multiple developers engaged in rapid communication and diverse conversation topics, the resulting chat messages often manifest complexity and lack structure. To enhance the efficiency of extracting information from chat threads , automatic mining techniques are introduced for thread classification. However, previous approaches still grapple with unsatisfactory classification accuracy due to two primary challenges that they struggle to adequately capture long-distance dependencies within chat threads and address the issue of category imbalance in labeled datasets. To surmount these challenges, we present a topic classification approach for chat information types named EAEChat. Specifically, EAEChat comprises three core components: the text feature encoding component captures contextual text features using a multi-head self-attention mechanism-based text feature encoder, and a siamese network is employed to mitigate overfitting caused by limited data; the data augmentation component expands a small number of categories in the training dataset using a technique tailored to developer chat messages, effectively tackling the challenge of imbalanced category distribution; the non-text feature encoding component employs a feature fusion model to integrate deep text features with manually extracted non-text features. Evaluation across three real-world projects demonstrates that EAEChat, respectively, achieves an average precision, recall, and F1-score of 0.653, 0.651, and 0.644, and it marks a significant 7.60% improvement over the state-of-the-art approaches. These findings confirm the effectiveness of our method in proficiently classifying developer chat messages in online chatrooms. Xiuwei Shang, Shikai Guo, Yulong Li 0001, Rong Chen 0003, Hui Li 0014, He Jiang 0001 |
ACM Trans. Softw. Eng. Methodol. | 7 |
| 2024 | Code Comment Inconsistency Detection Based on Confidence LearningabstractCode comments are a crucial source of software documentation that captures various aspects of the code. Such comments play a vital role in understanding the source code and facilitating communication between developers. However, with the iterative release of software, software projects become larger and more complex, leading to a corresponding increase in issues such as mismatched, incomplete, or outdated code comments. These inconsistencies in code comments can misguide developers and result in potential bugs, and there has been a steady rise in reports of such inconsistencies over time. Despite numerous methods being proposed for detecting code comment inconsistencies, their learning effect remains limited due to a lack of consideration for issues such as characterization noise and labeling errors in datasets. To overcome these limitations, we propose a novel approach called MCCL that first removes noise from the dataset and then detects inconsistent code comments in a timely manner, thereby enhancing the model's learning ability. Our proposed model facilitates better matching between code and comments, leading to improved development of software engineering projects. MCCL comprises two components, namely method comment detection and confidence learning denoising. The method comment detection component captures the intricate relationships between code and comments by learning their syntactic and semantic structures. It correlates the code and comments through an attention mechanism to identify how changes in the code affect the comments. Furthermore, confidence learning denoising component of MCCL identifies and removes characterization noises and labeling errors to enhance the quality of the datasets. This is achieved by implementing principles such as pruning noisy data, counting with probabilistic thresholds to estimate noise, and ranking examples to train with confidence. By effectively eliminating noise from the dataset, our model is able to more accurately learn inconsistencies between comments and source code. Our experiments on 1,518 open-source projects demonstrate that MCCL can accurately detect inconsistencies, achieving an averageF1-scoreof 82.6%. This result outperforms state-of-the-art methods by 2.4% to 28.0%. Therefore, MCCL is more effective in identifying inconsistent comments based on code changes compared to existing approaches. Zhengkang Xu, Shikai Guo, Rong Chen 0003, Hui Li 0014, He Jiang 0001 |
IEEE Trans. Software Eng. | 5 |
| 2023 | Constructing meaningful code changes via graph transformerabstractAbstract The rapid development of Open‐Source Software (OSS) has resulted in a significant demand for code changes to maintain OSS. Symptoms of poor design and implementation choices in code changes often occur, thus heavily hindering code reviewers to verify correctness and soundness of code changes. Researchers have investigated how to learn meaningful code changes to assist developers in anticipating changes that code reviewers may suggest for the submitted code. However, there are two main limitations to be addressed, including the limitation of long‐range dependencies of the source code and the missing syntactic structural information of the source code. To solve these limitations, a novel method is proposed, named Graph Transformer for learning meaningful Code Transformations (GTCT), to provide developers with preliminary and quick feedback when developers submit code changes, which can improve the quality of code changes and improve the efficiency of code review. GTCT comprises two components: code graph embedding and code transformation learning. To address the missing syntactic structural information of the source code limitation, the code graph embedding component captures the types and patterns of code changes by encoding the source code into a code graph structure from the lexical and syntactic representations of the source code. Subsequently, the code transformation learning component uses the multi‐head attention mechanism and positional encoding mechanism to address the long‐range dependencies limitation. Extensive experiments are conducted to evaluate the performance of GTCT by both quantitative and qualitative analyses. For the quantitative analysis, GTCT relatively outperforms the baseline on six datasets by 210%, 342.86%, 135%, 29.41%, 109.09%, and 91.67% in terms of perfect prediction. Meanwhile, the qualitative analysis shows that each type of code change by GTCT outperforms that of the baseline method in terms of bug fixed, refactoring code and others' taxonomy of code changes. Shikai Guo, Mengxuan Li 0005, Hui Li 0014, Rong Chen 0003 |
IET Softw. | 4 |
| 2023 | Structuring meaningful bug-fixing patches to fix software defectabstractAbstract Currently, software projects require a significant amount of time, effort and other resources to be invested in software testing to reduce the number of code defects. However, this process decreases the efficiency of software development and leads to a significant waste of workforce and resources. To address this challenge, researchers developed various solutions utilising deep neural networks. However, these solutions are frequently challenged by issues, such as a vast vocabulary, network training difficulties and elongated training processes resulting from the handling of redundant information. To overcome these limitations, the authors proposed a new neural network‐based model named HopFix, designed to detect software defects that may be introduced during the coding process. HopFix consists of four parts: data preprocessing, encoder, decoder and code generation components, which were used for preprocessing data, extracting information about software defects, analysing defect information, generating software patches and controlling the generation process of software patches, respectively. Experimental studies on Bug‐Fix Pairs (BFP) show that HopFix correctly fixed 47.2% ( BFP small datasets) and 25.7% ( BFP medium datasets) of software defects. Hui Li 0014, Xuexin Qi, Shikai Guo |
IET Softw. | 1 |
| 2023 | An data augmentation method for source code summarization
Zixuan Song, Xiuwei Shang, Guanxi Li, Hui Li 0014, Shikai Guo |
Neurocomputing | 5 |
| 2023 | Multi-Feature Fusion Based Structural Deep Neural Network for Predicting Answer Time on Stack Overflow
Shikai Guo, Hui Li 0014, Yu-Long Fan |
J. Comput. Sci. Technol. | 3 |
| 2023 | Code samples summarization for knowledge exchange in developer communityabstractAbstract A question title's function is to generate readable titles and describe a problem encountered by the code. Previous studies often used an end‐to‐end sequence‐to‐sequence system to generate question title's from source code. However, long‐term dependencies are often difficult to capture, and this may result in an incomplete source code representation. To address this issue, we propose a Transformer for Generating Code Title (hereinafter referred to as TGCT) model. Specifically, the TGCT model uses the position coding mechanism to model paired relationships between source terms by applying relative position representations. Multiple self‐attention mechanism components are also used to capture long‐term dependencies of the code. Comprehensive experiments on datasets from five coding languages, namely Python, Java, JavaScript, C#, and SQL, are conducted, and the results show that TGCT outperforms state‐of‐the‐art models based on the measurements of BLEU and ROUGE in general. In addition, a cross‐sectional comparison experiment was conducted to verify the effects of different model parameters, different data set sizes, position coding mechanism, and self‐attention mechanism on model results. Shikai Guo, Zhongyan Liu, Zixuan Song, Hui Li 0014, Rong Chen 0003 |
Softw. Pract. Exp. | 4 |
| 2023 | Feature transfer learning by reinforcement learning for detecting software defectabstractAbstract Software defects, produced inevitably in software projects, seriously affect the efficiency of software testing and maintenance. An appealing solution is the software defect prediction (SDP) that has achieved good performance in many software projects. However, the difference between features and the difference of the same feature between training data and test data may degrade defect prediction performance if such differences violate the model's assumption. To address this issue, we propose a SDP method based on feature transfer learning (FTL), which performs a transformation sequence for each feature in order to map the original features to another feature space. Specifically, FTL first uses the reinforcement learning scheme that automatically learns a strategy for transferring the potential feature knowledge from the training data. Then, we use the learned feature knowledge to inspire the transformation of the test data. The classifier is trained by the transformed training data and predicts defects for transformed test data. We evaluate the validity of FTL on 43 projects from PROMISE and NASA MDP using three classifiers, logistic regression, random forest, and Naive Bayes (NB). Experimental results indicate that FTL is better than the original classifiers and has the best performance on the NB classifier. For PROMISE, after using FTL, the average results of F1‐score, AUC, MCC are 0.601, 0.757, and 0.350 respectively, which are 24.9%, 2.6%, and 16.7% higher than the original NB classifier results. The number of projects with improved performance accounts for 83.87%, 83.87%, and 64.52%. Similarly, FTL performs well on NASA MDP. Besides, compared with four feature engineering (FE) methods, FTL achieves an excellent improvement on most projects and the average performance is also better than or close to the FE methods. Shikai Guo, Hui Li 0014, Rong Chen 0003 |
Softw. Pract. Exp. | 5 |
| 2023 | An Easy Data Augmentation Approach for Application Reviews Event InferenceabstractApplication review event inference aims to assess the effectiveness of application problems in response to user actions, which enables application developers to promptly discover and address potential issues in various applications, thereby improving their development and maintenance efficiency. Despite the development of event inference models for app reviews, which extract them as user action and app problem events and establish a relationship model between events and inference labels, the accuracy of these models is constrained due to limitations in labeling and characterizing noise and the lack of robustness and generalization. To address this challenge, we propose a model called Easy Data Augmentation for Application Reviews Event Inference (short for EDA-AREI), which comprises a denoising component, data augmentation component, and event inference prediction component. Specifically, the denoising component identifies labels and characterizes noisy data to enhance dataset quality, the data augmentation component replaces non-stop words with synonyms to increase textual diversity, and the event inference and prediction component reconstructs the classifier using denoised and augmented data. Experimental results on six datasets of one-star app reviews in the Apple App Store demonstrate that the EDA-AREI method achieves anAccuracyof 71.19%, 79.14%, 69.05%, 69.02%, 68.24% and 68.48%, respectively, representing an improvement of 0.83%–2.09% compared to state-of-the-art models. Regarding theF1-score, EDA-AREI achieves values of 71.30%, 69.93%, and 68.76% on the threshold_0.5, k-means_2, and random datasets, respectively, outperforming state-of-the-art models by 1.89%–4.02%. Furthermore, EDA-AREI achievesAUCvalues of 75.66% and 73.37% on the threshold_0.5 and k-means_2 datasets, respectively. As a result, EDA-AREI demonstrates substantial improvements inAccuracy, as well as enhancedF1-scoreandAUCacross most datasets, thereby enhancing the model's accuracy and robustness in identifying related action-problem pairs. Shikai Guo, Haorui Lin, Jiaoru Zhao, Hui Li 0014, Rong Chen 0003, He Jiang 0001 |
IEEE Trans. Software Eng. | 4 |
| 2023 | DupHunter: Detecting Duplicate Pull Requests in Fork-Based DevelopmentabstractThe emergence of numerous fork-based development platforms facilitates the development of Open-Source Software (OSS) projects. Developers across the world can fork software projects and submit their Pull Requests (PRs) to the projects. However, as the number of forks increases, numerous duplicate PRs might be submitted. These duplicate PRs may cause extra code review workload and frustrate developers working on the projects. To detect duplicate PRs, many approaches have been proposed, which analyze the similarity of different elements in PRs. However, previous approaches still suffer from unsatisfied detection accuracy due to two challenges. That is, they ignore the syntactic structural information of text elements in PRs and lack the joint reasoning between different elements of two PRs. In this study, we propose an automated duplicate PRs detector namedDupHunter(Duplicate PRsHunter), which includes a graph embedding component and a duplicate PRs detection component to address the above challenges. The graph embedding component uses a feature graph to represent a PR. It encodes the syntactic structure and semantics of text elements (e.g., the title and the description), as well as the knowledge of non-text elements (e.g., the submission time), to address the syntactic structural information challenge. The duplicate PRs detection component tackles the joint reasoning challenge using a graph matching network, which enables the information exchange and matching across different elements of two feature graphs with an attention coefficient mechanism. Experiments on 26 open-source projects show that DupHunter achieves an averageF1-score@1value of 0.650, significantly outperforming the state-of-the-art approaches by 3.2% to 48.1%. DupHunter can accurately detect duplicate PRs, with an averagePrecision@1value of 0.922 and an averageRecall@1value of 0.502. He Jiang 0001, Yulong Li 0001, Shikai Guo, Tao Zhang 0001, Hui Li 0014, Rong Chen 0003 |
IEEE Trans. Software Eng. | 6 |
| 2022 | Identifying High-impact Bug Reports with Imbalance Distribution by Instance Fuzzy EntropyabstractBug tracking systems, such as Bugzilla, contain bug reports collected from sources such as development teams, testing teams and end users. Developers often depend on bug reports to fix identified bugs. Frequently used bug reports are the so-called severe bug reports. Although severe bug reports can be manually detected within bug reports in bug tracking systems, they impose heavy burdens on management of bug tracking systems. Consequently, an automated mechanism to examine the severity of bug reports is desirable to augment productivity. Unfortunately, identifying the severity of bug reports from thousands of bug reports in a bug tracking system is not an easy feat, because of the problem of low-quality and imbalance distributions that could affect the performance of automated mechanisms. In this paper, we propose an approach, namely FER, to counter low-quality and imbalanced distributions of bug reports relative to their severity. First, FER approach gets high-quality bug reports based on instance fuzzy entropy. Then, FER approach weakens the imbalancedness degree of class distribution according to the high-quality bug reports to train classifiers to recognize the severity of bug reports. Several experiments are conducted on bug reports from three open source projects (Eclipse, Mozilla, GNOME) and they reveal that our approach is robust against the low-quality and imbalance distributions of bug reports, while identifying the severity of bug reports. Hui Li 0014, Xuexin Qi, Mengxuan Li 0005 |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2022 | Self-admitted technical debt detection by learning its comprehensive semantics via graph neural networksabstractAbstract The goal of software development is to deliver software products with high quality and free from defects, but resource and time constraints often cause the developers to submit incomplete or temporary patches of codes and further bear the additional burden. Therefore, the investigations on identifying self‐admitted technical debt (SATD) to improve code quality have been conducted in recent years. However, missing syntactic structure information and the imbalance distribution bias shorten the SATD identification performance. Addressing to this issue, we present a graph neural network based SATD identification model (GNNSI) to improve the performance. Specifically, we obtain the structure information of the missing SATD in a compositional way to obtain different feature maps for different comments, and use focal loss to handle the imbalance between SATD and non‐SATD classes in the comments. Then extensive experiments on 10 open source projects are conducted, and the results show that GNNSI outperforms the baselines and can help developers to better predict SATDs. Hui Li 0014, Rong Chen 0003, Jun Ai, Shikai Guo |
Softw. Pract. Exp. | 1 |
| 2022 | Quantized Output-Feedback Control for Unmanned Marine Vehicles With Thruster Faults via Sliding-Mode TechniqueabstractThis article is concerned with the quantized output-feedback control problem for unmanned marine vehicles (UMVs) with thruster faults and ocean environment disturbances via a sliding-mode technique. First, based on output information and compensator states, an augmented sliding surface is constructed and sliding-mode stability through linear matrix inequalities can be guaranteed. An improved quantization parameter dynamic adjustment scheme, with a larger quantization parameter adjustment range, is then given to compensate for quantization errors effectively. Combining the quantization parameter adjustment strategy and adaptive mechanism, a novel robust sliding-mode controller is designed to guarantee the asymptotic stability of a closed-loop UMV system. As a result, a smaller lower bound of the thruster fault factor than that of the existing result can be tolerated, which brings more practical applications. Finally, the comparison simulation results have illustrated the effectiveness of the proposed method. Tieshan Li 0001, Hui Li 0014 |
IEEE Trans. Cybern. | 4 |
| 2021 | Software defect prediction with imbalanced distribution by radius-synthetic minority over-sampling techniqueabstractAbstract Software defect prediction, which can identify the defect‐prone modules, is an effective technology to ensure the quality of software products. Due to the importance in software maintenance, many learning‐based software defect prediction models are presented in recent years. Actually, the defects usually occupy a very small proportions in software source codes; thus, the imbalanced distributions between defect‐prone modules and non‐defect‐prone modules increase the learning difficulty of the classification task. To address this issue, we present a random over‐sampling mechanism used to generate minority‐class samples from high‐dimensional sampling space to deal with the imbalanced distributions in software defect prediction, in which two constraints are applied to provide a robust way to generate new synthetic samples, that is, scaling the random over‐sampling scope to a reasonable area and distinguishing the majority‐class samples in a critical region. Based on nine open datasets of software projects, we experimentally verify that our presented method is effective on predict the defect‐prone modules, and the effect is superior to the traditional imbalanced processing methods. Shikai Guo, Hui Li 0014 |
J. Softw. Evol. Process. | 3 |
| 2021 | Underwater image enhancement with image colorfulness measure
Xi Yang 0009, Hui Li 0014, Rong Chen 0003 |
Signal Process. Image Commun. | 2 |
| 2021 | Quantized Sliding Mode Control of Unmanned Marine Vehicles: Various Thruster Faults Tolerated With a Unified ModelabstractThis paper investigates quantized sliding mode control of unmanned marine vehicles (UMVs) with thruster faults and nonlinearities. We give a unified model to accommodate different types of thruster faults (e.g., partial, total, time-varying stuck, hard-over, and bias faults) in a common framework, which is significant because existing methods can only address them separately in a fault-specific manner. To eliminate the quantization effect induced by the communication channel by which the UMV outputs (e.g., position and velocity) and the control inputs are transmitted to and from the remote control station, a new dynamic uniform quantizer with an adjustable range of sensitivity is given. Via flexible choice of parameters, the adjustment range can fall within that of the existing results in the fault-free case. A quantized sliding mode controller and a dynamic quantization parameter adjustment strategy are then developed to suppress oscillation amplitudes of the yaw velocity error and the yaw angle in the presence of thruster faults. Simulation studies have verified the effectiveness of the proposed method. Ge Guo 0001, Hui Li 0014 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2020 | Developer Activity Motivated Bug Triaging: Via Convolutional Neural Network
Shikai Guo, Xi Yang 0009, Rong Chen 0003, Chen Guo 0001, Hui Li 0014 |
Neural Process. Lett. | 6 |
| 2019 | Identify Severity Bug Report with Distribution Imbalance by CR-SMOTE and ELMabstractManually inspecting bugs to determine their severity is often an enormous but essential software development task, especially when many participants generate a large number of bug reports in a crowdsourced software testing context. Therefore, boosting the capabilities of methods of predicting bug report severity is critically important for determining the priority of fixing bugs. However, typical classification techniques may be adversely affected when the severity distribution of the bug reports is imbalanced, leading to performance degradation in a crowdsourcing environment. In this study, we propose an enhanced oversampling approach called CR-SMOTE to enhance the classification of bug reports with a realistically imbalanced severity distribution. The main idea is to interpolate new instances into the minority category that are near the center of existing samples in that category. Then, we use an extreme learning machine (ELM) — a feedforward neural network with a single layer of hidden nodes — to predict the bug severity. Several experiments were conducted on three datasets from real bug repositories, and the results statistically indicate that the presented approach is robust against real data imbalance when predicting the severity of bug reports. The average accuracies achieved by the ELM in predicting the severity of Eclipse, Mozilla, and GNOME bug reports were 0.780, 0.871, and 0.861, which are higher than those of classifiers by 4.36%, 6.73%, and 2.71%, respectively. Shikai Guo, Rong Chen 0003, Hui Li 0014, Tianlun Zhang |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2019 | The Influence Ranking for Testers in Bug Tracking SystemsabstractAt present, bug tracking systems are used to collect and manage bug reports in many software projects. As participants, the testers not only submit bug reports to the system, but also comment on bug reports in the system. The tester’s behaviors of submitting and commenting reflect his/her influence in bug tracking systems. However, with the rapid increase of the bug reports in software projects, evaluating the testers’ influence in the projects accurately becomes more and more difficult. Aiming at solving this problem, the submission and comment on bug report can be regarded as social behaviors of the testers, and thus the method of Influence Ranking for Testers (IRfT) in bug tracking systems is presented and used for measuring the influence of the testers in this paper. The case study of the Eclipse project in Bugzilla shows that the result produced by IRfT is consistent with the actual performance of the testers in this project. The ranking results can keep stable in the cases of link adding or removing and tester removing in tester networks, and the results are also proved to be valid in the future. The further investigation on the speed of network break-down by node removal demonstrates that the top-ranking testers are important in the organization of tester networks. Additionally, the results also show that the ranking of the testers is related to the existence time in bug tracking system. Therefore, IRfT is proved to be an effective measurement for evaluating the influence of the testers in bug tracking system, and it can further demonstrate the testers’ contributions in software testing, such as bug validations, bug fixes, etc. Hui Li 0014, Guofeng Gao, Rong Chen 0003, Shikai Guo |
Int. J. Softw. Eng. Knowl. Eng. | 1 |
| 2019 | Single Image Haze Removal via Region Detection NetworkabstractHaze removal typically works on a physical model to estimate how light is transmitted and lost due to absorption and scattering through the atmosphere. In this paper, a region detection network is proposed to learn the relationship between the hazy image and the medium transmission map in a patchwise manner; the transmission map is then used to remove haze via an atmospheric scattering model and enhance the detail of de-hazed images. To this end, we design a simple yet powerful deep convolutional neural network, which mainly consists of two types of network units and can be trained in an end-to-end manner. One network unit is a module with the residual structure that facilitates the learning process of the deep network. The other is a novel module with a cascaded cross channel pool, which fuses multi-level haze-relevant features and boosts the abstraction ability of the model on a nonlinear manifold. Moreover, an evolutionary-based enhancement method is developed to improve the level of detail of over-smoothed results. Several comparative experiments have been conducted on synthetic and real images, through which we conclude that the proposed method achieves state-of-the-art haze removal results, qualitatively and quantitatively. Supplementary experiments further indicate that our method works better against other adverse effects on vision quality (e.g., the mist formed by heavy rain and the veil met underwater). Moreover, we present a lightweight version of the proposed network, which achieves an impressive haze removal performance even on low-power devices. Xi Yang 0009, Hui Li 0014, Yu-Long Fan, Rong Chen 0003 |
IEEE Trans. Multim. | 2 |
| 2018 | Capability Matching and Heuristic Search for Job Assignment in Crowdsourced Web Application TestingabstractWeb based commercial systems are increasingly becoming feature rich, interactive and functional as locally installed applications. Testing web applications is unique, as many factors affect the system performance and user experience. Crowdsourcing is an appealing and economic solution to web application testing due to the ability to reach a larger international audience. However, less is known about the quality control of crowdsourced testing to harness the collective efforts of individuals. In our study, the collaborative testing problem in a crowdsourcing environment is defined as a job assignment problem and is formulated as an integer linear programming (ILP) problem. The objective of this paper is to validate a greedy job assignment approach as a tool for the effective use of crowdsourced testing. We carried out a case study on Xturk, a prototype crowdsourced testing system, to understand the crowdsourced testers behaviour that the trustworthiness, the execution time of test cases and accuracy of feedback. Several experiments indicate that this approach is comparatively effective with regards to the feasibility verdict, efficiency and accuracy. Shikai Guo, Rong Chen 0003, Hui Li 0014 |
SMC | 3 |
| 2018 | Crowdsourced Web Application Testing Under Real-Time ConstraintsabstractCrowdsourcing carried out by cyber citizens instead of hired consultants and professionals has become increasingly an appealing solution to test the feature rich and interactive web. Despite having various online crowdsourcing testing services, the benefits of exposure to a wider audience and harnessing the collective efforts of individuals remain uncertain, especially when the quality control is problematic in an open environment. The objective of this paper is to propose a real-time collaborative testing approach (RCTA) to create a productive crowdsourced testing on a dynamic Internet. We implemented a prototype crowdsourcing system XTurk, and carried out a case study, to understand the crowdsourced testers behavior, the trustworthiness, the execution time of test cases and accuracy of feedback. Several experiments are carried out and experimental results validate the quality, efficiency and reliability of the present approach and the positive testing feedback is are shown to outperform the previous methods. Shikai Guo, Rong Chen 0003, Hui Li 0014, Jian Gao 0007 |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2014 | Extraction and Analysis of Crucial Fraction in Software NetworksabstractMany complex systems, such as software systems, are full of complexity arising from interactions among basic units (such as classes, interfaces and struts in object-oriented software systems). One of the most successful approaches to capture the underlying structural features of large-scale software systems is the investigation of hierarchical organization. However, the hierarchy of software networks has not been thoroughly investigated. In this paper, the crucial fraction (CF) in software networks has been extracted and analyzed in a set of real-world software systems. First, the classes and the relationships between them have been extracted into software networks. Then software networks have been divided into different layers, and CF of software networks has been extracted by k-core. The empirical studies in this paper reveal that software networks represent flat hierarchical structure. Finally, CF has been measured by the relevant complex network parameters respectively, and the relations between CF and overall network have been analyzed by the case studies of software networks. The results show that CF represents characteristics of scale-free, small-world, strong connectivity, and the units in CF are frequently reused and dominate the overall system. Hui Li 0014, Rong Chen 0003, Hai Zhao 0002 |
Int. J. Softw. Eng. Knowl. Eng. | 1 |