EDBT 2026 Demo / reviewers in the wild / expert
Zhao Wei
dblp:77/1533
· DBLP profile ↗
26ranked-venue papers
4as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 11 · 11 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSecurity and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Landscape-aware Automated Algorithm Design: An Efficient Framework for Real-world OptimizationabstractThe advent of Large Language Models (LLMs) has opened new frontiers in automated algorithm design, giving rise to numerous powerful methods. However, these approaches retain critical limitations: they require extensive evaluation of the target problem to guide the search process, making them impractical for real-world optimization tasks, where each evaluation consumes substantial computational resources. This research proposes an innovative and efficient framework that decouples algorithm discovery from high-cost evaluation. Our core innovation lies in combining a Genetic Programming (GP) function generator with an LLM-driven evolutionary algorithm designer. The evolutionary direction of the GP-based function generator is guided by the similarity between the landscape characteristics of generated proxy functions and those of real-world problems, ensuring that algorithms discovered via proxy functions exhibit comparable performance on real-world problems. Our method enables deep exploration of the algorithmic space before final validation while avoiding costly real-world evaluations. We validate the framework's efficacy across multiple real-world problems, demonstrating its ability to discover high-performance algorithms while substantially reducing expensive evaluations. This approach shows a path to apply LLM-based automated algorithm design to computationally intensive real-world optimization challenges. Haoran Yin 0003, Shuaiqun Pan, Zhao Wei, Jian Cheng Wong, Yew-Soon Ong, Anna V. Kononova, Thomas Bäck, Niki van Stein |
GECCO | 3 |
| 2025 | Multi-Head Auto-Correlation Attention Networks for Session-based Social RecommendationabstractSession-based Social Recommendation (SSR) aims to improve next-item prediction by combining a user’s session activities with insights from their social networks. However, the brevity of sessions makes SSR models prone to noise, and many methods rely on complex Deep Neural Networks (DNNs), which often lead to long training times. To address these challenges, we introduce a novel SSR approach that leverages autocorrelation, a concept from stochastic processes, to model sequential dependencies more effectively. By applying Fast Fourier Transforms (FFT) to compute autocorrelation and integrating this with multi-head attention mechanisms, our method captures temporal patterns within session data more accurately. Extensive experiments on two public datasets demonstrate that our method outperforms existing state-of-the-art SSR models in both accuracy and efficiency. Our code is available at https://github.com/LUUUUUUZ/MACNet. Mengying Lu, Zhao Wei, Bingxu An |
ICASSP | 4 |
| 2025 | Evolvable Conditional DiffusionabstractThis paper presents an evolvable conditional diffusion method such that black-box, non-differentiable multi-physics models, as are common in domains like computational fluid dynamics and electromagnetics, can be effectively used for guiding the generative process to facilitate autonomous scientific discovery. We formulate the guidance as an optimization problem where one optimizes for a desired fitness function through updates to the descriptive statistic for the denoising distribution, and derive an evolution-guided approach from first principles through the lens of probabilistic evolution. Interestingly, the final derived update algorithm is analogous to the update as per common gradient-based guided diffusion models, but without ever having to compute any derivatives. We validate our proposed evolvable diffusion algorithm in two AI for Science scenarios: the automated design of fluidic topology and meta-surface. Results demonstrate that this method effectively generates designs that better satisfy specific optimization objectives without reliance on differentiable proxies, providing an effective means of guidance-based diffusion that can capitalize on the wealth of black-box, non-differentiable multi-physics numerical models common across Science. Zhao Wei, Chin Chun Ooi, Abhishek Gupta 0001, Jian Cheng Wong, Pao-Hsiung Chiu, Sheares Xue Wen Toh, Yew-Soon Ong |
IJCAI | 1 |
| 2025 | SmellDetector: Multi-Label Code Smell Detection and Refactoring with Large Language ModelsabstractLarge Language Models (LLMs) have demonstrated impressive capabilities in many tasks such as code generation and automated program repair. However, code LLMs have ignored another important task in programmers’ daily development work, which is to improve the maintainability, readability, and scalability of the program. All of these characteristics are related to code smells and we study how to improve them by detecting and removing code smells. Most works on code smells still rely on using measures formulated by experts as features, but lack of use of the rich prior knowledge contained in code LLMs. In this paper, we propose SmellDetector, a comprehensive model for both code smell detection and refactoring opportunities detection in Java. We train the model with the designed prompt which contains both code smells of class-level and method-level in the same code snippet, including more than 20 types. We achieve state-of-the-art performance on the code smell detection task and change the basic paradigm of code smell detection from binary classification problem to multi-label classification. Finally, it has been verified through experiments that good code smell detection helps to detect refactoring opportunities. Hai-Tao Zheng 0002, Haiye Lin, Hong-Gee Kim, Bingxu An, Zhao Wei, Yong Xu 0007 |
IJCNN | 8 |
| 2025 | Repository-Level Code Smell Detection Based on Multi-Scale Code Information and LLM AssistanceabstractIn software engineering, code smell detection has always been an important research task because it affects the readability and maintainability of programs. Traditional code smell detection work mostly focuses on a single file, and the study of repo-level code smell caused by the interaction of multiple files is still relatively lacking, although it is more practical in reality. Although large language models (LLMs) have recently achieved remarkable success in the field of code generation, simply copying the fine-tuning method based on LLM is not good enough in this task because it is difficult to model the logical relationship between multiple files. In this paper, we propose a repo-level code smell detection model(RSD), which divides the levels according to the distance between the rest class and the problem class, uses cross attention and CNN to model global and local information, and uses LLM to generate teacher vectors in the same scenario for assistance. Finally, the detection performance surpasses the fine-tuning baseline method based on multiple code LLMs and has achieved the state-of-the-art. Finally, we also contribute the BenchMark used in this article to help researchers in the repo-level smell detection task. Yongqin Zeng, Hai-Tao Zheng 0002, Haiye Lin, Hong-Gee Kim, Bingxu An, Zhao Wei, Yong Xu 0007 |
IJCNN | 8 |
| 2025 | DLCoG: A Novel Framework for Dual-Level Code Comment Generation Based on Semantic Segmentation and In-Context LearningabstractIn large software projects with collaborative development, comprehensive code comments are crucial for code readability and maintainability. Code comments mainly include method comments and inline comments, where the former describes the functionality globally, and the latter describes the implementation details locally. Existing methods typically generate these two kinds of comments with specific locations independently, which results in weak correlations between comments and code context, as well as high model inference costs due to long token inputs. To address these issues, we define the combination of inline comments and method comments as Dual-Level Code Comments. We formulate the novel task of automatically generate dual-level code comments based on given code and propose an approach named DLCoG (DualLevel Code Comment Generation) to automate this task. First, a Semantic Segmentation and Identification multi-task model based on CodeBERT, termed Se2Iden (Semantic Segmentation and Identification model), is proposed to identify code segments requiring inline comments. Next, we retrieve similar samples to adopting the in-context learning paradigm, which can enhance the generation quality of large language models (LLMs) in specific domains. Finally, the LLM is guided to generate duallevel code comments using Chain-of-Thought (CoT) prompts that first produce inline comments, followed by method comments. We manually constructed a high-quality clean Java dataset consisting of*> based on open-source Java projects by (i) determining comments type and (ii) manually associating inline comments with their corresponding code. Then, we trained a multi-task learning model based on CodeBERT to automatically take the two steps needed, termed ICSA (Inline Comment Classification and Scope Association), thus to expand to a dataset containing 80k dual-level code comments. Experimental results on clean and extended datasets show that DLCoG outperforms all baselines by substantial margins. The contextual information provided by DLCoG can effectively improve the inline comments generated by LLM. Coordinated generation of dual-level comment also brings effective improvements to method comments, which is particularly significant when there are few contextual examples. Our work fills the long-standing gap in the dual-level code comment generation field, and can provide insights for future research in this direction. We provide open-source datasets and source code for future research. Haiyang Yang, Qingyang Yan, Weihuan Min, Zhao Wei, Li Kuang, Yingjie Xia |
ICPC | 6 |
| 2025 | MCTS-Refined CoT: High-Quality Fine-Tuning Data for LLM-Based Repository Issue ResolutionabstractLLMs demonstrate strong performance in automated software engineering, particularly for code generation and issue resolution. While proprietary models like GPT-4o achieve high benchmarks scores on SWE-bench, their API dependence, cost, and privacy concerns limit adoption. Open-source alternatives offer transparency but underperform in complex tasks, especially sub-100B parameter models. Although quality Chain-of-Thought (CoT) data can enhance reasoning, current methods face two critical flaws: (1) weak rejection sampling reduces data quality, and (2) inadequate step validation causes error accumulation. These limitations lead to flawed reasoning chains that impair LLMs’ ability to learn reliable issue resolution.The paper proposes MCTS-REFINE, an enhanced Monte Carlo Tree Search (MCTS)-based algorithm that dynamically validates and optimizes intermediate reasoning steps through a rigorous rejection sampling strategy, generating high-quality CoT data to improve LLM performance in issue resolution tasks. Key innovations include: (1) augmenting MCTS with a reflection mechanism that corrects errors via rejection sampling and refinement, (2) decomposing issue resolution into three subtasks—File Localization, Fault Localization, and Patch Generation—each with clear ground-truth criteria, and (3) enforcing a strict sampling protocol where intermediate outputs must exactly match verified developer patches, ensuring correctness across reasoning paths.Experiments on SWE-bench Lite and SWE-bench Verified demonstrate that LLMs fine-tuned with our CoT dataset achieve substantial improvements over baselines. Notably, Qwen2.5-72B-Instruct achieves 28.3%(Lite) and 35.0%(Verified) resolution rates, surpassing SOTA baseline SWE-Fixer-Qwen-72B with the same parameter scale, which only reached 24.7%(Lite) and 32.8%(Verified). Given precise issue locations as input, our fine-tuned Qwen2.5-72B-Instruct model achieves an impressive issue resolution rate of 43.8%(Verified), comparable to the performance of Deepseek-v3. We open-source our MCTS-REFINE framework, CoT dataset, and fine-tuned models to advance research in AI-driven software engineering. Yibo Wang 0008, Zhihao Peng 0009, Ying Wang 0038, Zhao Wei, Hai Yu 0001, Zhiliang Zhu 0001 |
ASE | 4 |
| 2025 | Deep Code Search with Naming-Agnostic Contrastive Multi-View LearningabstractSoftware development is a repetitive task, as developers usually reuse or get inspiration from existing implementations. Code search, which refers to the retrieval of relevant code snippets from a codebase according to the developer’s intent that has been expressed as a query, has become increasingly important in the software development process. Due to the success of deep learning in various applications, a great number of deep learning-based code search approaches have sprung up and achieved promising results. However, developers may not follow the same naming conventions and the same variable may have different variable names in different implementations, bringing a challenge to deep learning-based code search methods that rely on explicit variable correspondences to understand source code. To overcome this challenge, we propose a Naming-Agnostic Code Search (NACS) method based on contrastive multi-view code representation learning. NACS strips information bound to variable names from Abstract Syntax Tree (AST), the representation of the abstract syntactic structure of source code, and focuses on capturing intrinsic properties solely from AST structures. We use semantic-level and syntax-level augmentation techniques to prepare realistically rational data and adopt contrastive learning to design a graph-view modeling component in NACS to enhance the understanding of code snippets. We further model ASTs in a path view to strengthen the graph-view modeling component through multi-view learning. Extensive experiments show that NACS provides superior code search performance compared to baselines and NACS can be adapted to help existing code search methods overcome the impact of different naming conventions. Our implementation is available at https://github.com/KDEGroup/NACS . Jiadong Feng, Wei Li 0274, Suhuang Wu, Zhao Wei, Yong Xu 0007, Juhong Wang, Hui Li 0057 |
ACM Trans. Knowl. Discov. Data | 4 |
| 2025 | On the Effectiveness of Large Language Models in Domain-Specific Code GenerationabstractLarge language models (LLMs) such as ChatGPT have shown remarkable capabilities in code generation. Despite significant achievements, they rely on enormous training data to acquire a broad spectrum of open-domain knowledge. Besides, their evaluation revolves around open-domain benchmarks like HumanEval, which primarily consist of programming contests. Therefore, it is hard to fully characterize the intricacies and challenges associated with particular domains (e.g., Web, game, and math). In this article, we conduct an in-depth study of the LLMs in domain-specific code generation. Our results demonstrate that LLMs exhibit sub-optimal performance in generating domain-specific code, due to their limited proficiency in utilizing domain-specific libraries. We further observe that incorporating API knowledge as prompts can empower LLMs to generate more professional code. Based on these findings, we further investigate how to effectively incorporate API knowledge into the code generation process. We experiment with three strategies for incorporating domain knowledge, namely, external knowledge inquirer, chain-of-thought prompting, and chain-of-thought fine-tuning. We refer to these strategies as a new code generation approach called DomCoder . Experimental results show that all strategies of DomCoder improve the effectiveness of domain-specific code generation under certain settings. Xiaodong Gu 0002, Yalan Lin, Hongyu Zhang 0002, Chengcheng Wan 0001, Zhao Wei, Yong Xu 0010, Juhong Wang |
ACM Trans. Softw. Eng. Methodol. | 7 |
| 2025 | Fine-Tuning Large Language Models to Improve Accuracy and Comprehensibility of Automated Code ReviewabstractAs code review is a tedious and costly software quality practice, researchers have proposed several machine learning-based methods to automate the process. The primary focus has been on accuracy, that is, how accurately the algorithms are able to detect issues in the code under review. However, human intervention still remains inevitable since results produced by automated code review are not 100% correct. To assist human reviewers in making their final decisions on automatically generated review comments, the comprehensibility of the comments underpinned by accurate localization and relevant explanations for the detected issues with repair suggestions is paramount. However, this has largely been neglected in the existing research. Large language models (LLMs) have the potential to generate code review comments that are more readable and comprehensible by humans, thanks to their remarkable processing and reasoning capabilities. However, even mainstream LLMs perform poorly in detecting the presence of code issues because they have not been specifically trained for this binary classification task required in code review. In this article, we contribute Comprehensibility of Automated Code Review using Large Language Models ( Carllm ), a novel fine-tuned LLM that has the ability to improve not only the accuracy but, more importantly, the comprehensibility of automated code review, as compared to state-of-the-art pre-trained models and general LLMs. Yongda Yu, Guoping Rong, Haifeng Shen, He Zhang 0001, Dong Shao, Zhao Wei, Juhong Wang |
ACM Trans. Softw. Eng. Methodol. | 7 |
| 2024 | Importance of Nyquist-Shannon Sampling in Training of Physics-Informed Neural NetworksabstractRecent rapid advances in artificial intelligence methods and computational resources have opened up the possibility of solving complex problems in physics with machine learning. While promising, the training of performant deep learning models remains challenging, especially as the choice of various neural architectures or model hyper-parameters can all greatly impact model performance. Hence, it is now fairly common for researchers to systematically conduct either a neural architecture search (NAS) or hyper-parameter tuning as part of the model training process. However, these usually involve repeated model training and evaluation (typically via a cross-validation framework) and can be very computationally expensive. While there have been training-free heuristics proposed for NAS in recent years, this remains an active area of work, and much remains unresolved. Concurrently, physics-informed neural networks (PINNs) have been proposed in recent years as an attractive means of incorporating physics-derived knowledge into a deep learning framework for the modelling of real-world physical systems. However, recent literature has begun to emerge on the additional complexity of training PINNs due to the addition of these physics-derived equations (typically in the form of partial differential equations, PDEs). The extended optimization trajectories observed further exacerbate the aforementioned computational cost of NAS or hyper-parameter search. Hence, in this work, we propose a first demonstration of the utility of training-free heuristics for accelerating the selection of model parameters. In particular, we assess the correlation between basic concepts espoused in Nyquist-Shannon sampling, and the appropriate selection of mini-batch size for PINN training with stochastic gradient descent (specifically, the state-of-the-art ADAM algorithm). As a proof-of-concept, this paper focuses on training PINNs to predict the solution to a canonical 2D Helmholtz equation with varying complexity (as measured by the frequency of the signal present), and we evaluate and report the influence of different mini-batch size parameters on the effective training of PINNs. The results show that PINN model performance exhibits a step-like phase transition with increased mini-batch size during training. Critically, we note that the minimum mini-batch size used increases proportionally with the presence of higher frequencies in the physical system being modelled, as is intuitively described by the Nyquist-Shannon sampling theorem. This both shows the utility of similar potential heuristics for the training-free selection of hyper-parameters when training PINNs and suggests the potential for insights towards accelerated design and training of PINNs for physical systems from prior theoretical concepts in signal processing and reconstruction. Chin Chun Ooi, Anran Huang, Zhao Wei, Shiyao Qin, Jian Cheng Wong, Pao-Hsiung Chiu, My Ha Dao |
IJCNN | 3 |
| 2024 | A Just-in-time Software Defect Localization Method based on Code Graph RepresentationabstractTraditional software defect localization aims to locate defective files, methods, or code lines based on symptoms such as defect reports. In comparison, Just-In-Time (JIT) software defect localization focuses on identifying defective code lines when a defective code change is initially submitted. It can identify issues at the code line level before the defect becomes apparent, preventing it from adversely affecting the software. Although researchers have proposed various methods for JIT defect localization, existing methods still have the following shortcomings: (1) Most methods rely heavily on tokens from single code lines to calculate naturalness for defect localization, which makes it challenging to effectively distinguish between code lines that have the same content but different labels (defective code lines or non-defective code lines) - termed Duplicate Lines with Different Labels (DLDL). (2) Existing methods represent code in the form of sequences, neglecting the structural information of the code. Therefore, we propose a JIT defect localization method based on code graph representation. First, we construct code linelevel code graphs for code changes to distinguish DLDL explicitly. Next, to extract sequential and structural information from the code, we propose a code graph representation model with contrastive learning to generate graph feature vectors and node scores with rich semantics. Finally, we calculate the naturalness of code lines based on the graph feature vectors and node scores. Using this naturalness, we identify defective code lines. Experimental results show that our JIT defect localization method outperforms the state-of-the-art methods. Huan Zhang 0017, Weihuan Min, Zhao Wei, Li Kuang, Honghao Gao, Huaikou Miao |
ICPC | 3 |
| 2024 | Joint attention mechanism with dynamic kernel for yolov5 mobile wireless charging coil surface defect identification
Zhao Wei |
Multim. Tools Appl. | 1 |
| 2024 | Esale: Enhancing Code-Summary Alignment Learning for Source Code Summarizationabstract(Source) code summarization aims to automatically generate succinct natural language summaries for given code snippets. Such summaries play a significant role in promoting developers to understand and maintain code. Inspired by neural machine translation, deep learning-based code summarization techniques widely adopt an encoder-decoder framework, where the encoder transforms given code snippets into context vectors, and the decoder decodes context vectors into summaries. Recently, large-scale pre-trained models for source code (e.g., CodeBERT and UniXcoder) are equipped with encoders capable of producing general context vectors and have achieved substantial improvements on the code summarization task. However, although they are usually trained mainly on code-focused tasks and can capture general code features, they still fall short in capturing specific features that need to be summarized. In a nutshell, they fail to learn the alignment between code snippets and summaries (code-summary alignment for short). In this paper, we propose a novel approach to improve code summarization based on summary-focused tasks. Specifically, we exploit a multi-task learning paradigm to train the encoder on three summary-focused tasks to enhance its ability to learn code-summary alignment, including unidirectional language modeling (ULM), masked language modeling (MLM), and action word prediction (AWP). Unlike pre-trained models that mainly predict masked tokens in code snippets, we design ULM and MLM to predict masked words in summaries. Intuitively, predicting words based on given code snippets would help learn the code-summary alignment. In addition, existing work shows that AWP affects the prediction of the entire summary. Therefore, we further introduce the domain-specific task AWP to enhance the ability of the encoder to learn the alignment between action words and code snippets. We evaluate the effectiveness of our approach, calledEsale, by conducting extensive experiments on four datasets, including two widely used datasets JCSD and PCSD, a cross-project Java dataset CPJD, and a multilingual language dataset CodeSearchNet. Experimental results show thatEsalesignificantly outperforms state-of-the-art baselines in all three widely used metrics, including BLEU, METEOR, and ROUGE-L. Moreover, the human evaluation proves that the summaries generated byEsaleare more informative and closer to the ground-truth summaries. Chunrong Fang, Weisong Sun, Zhao Wei, Quanjun Zhang, Yudu You, Bin Luo 0003, Yang Liu 0003, Zhenyu Chen 0001 |
IEEE Trans. Software Eng. | 5 |
| 2024 | Distilling Quality Enhancing Comments From Code Reviews to Underpin Reviewer RecommendationabstractCode review is an important practice in software development. One of its main objectives is for the assurance of code quality. For this purpose, the efficacy of code review is subject to the credibility of reviewers, i.e., reviewers who have demonstrated strong evidence of previously making quality-enhancing comments are more credible than those who have not. Code reviewer recommendation (CRR) is designed to assist in recommending suitable reviewers for a specific objective and, in this context, assurance of code quality. Its performance is susceptible to the relevance of its training dataset to this objective, composed of all reviewers’ historical review comments, which, however, often contains a plethora of comments that are irrelevant to the enhancement of code quality. Furthermore, recommendation accuracy has been adopted as the sole metric to evaluate a recommender's performance, which is inadequate as it does not take reviewers’ relevant credibility into consideration. These two issues form the ground truth problem in CRR as they both originate from the relevance of dataset used to train and evaluate CRR algorithms. To tackle this problem, we first propose the concept of Quality-Enhancing Review Comments (QERC), which includes three types of comments - change-triggering inline comments, informative general comments, and approve-to-merge comments. We then devise a set of algorithms and procedures to obtain a distilled dataset by applyingQERCto the original dataset. We finally introduce a new metric – reviewer's credibility for quality enhancement (RCQE) – as a complementary metric to recommendation accuracy for evaluating the performance of recommenders. To validate the proposed QERC-based approach to CRR, we conduct empirical studies using real data from seven projects containing over 82K pull requests and 346K review comments. Results show that: (a)QERCcan effectively address the ground truth problem by distilling quality-enhancing comments from the dataset containing original code reviews, (b)QERCcan assist recommenders in finding highly credible reviewers at a slight cost of recommendation accuracy, and (c) even “wrong” recommendations using the distilled dataset are likely to be more credible than those using the original dataset. Guoping Rong, Yongda Yu, He Zhang 0001, Haifeng Shen, Dong Shao, Hongyu Kuang, Zhao Wei, Juhong Wang |
IEEE Trans. Software Eng. | 9 |
| 2023 | RefBERT: A Two-Stage Pre-trained Framework for Automatic Rename RefactoringabstractRefactoring is an indispensable practice of improving the quality and maintainability of source code in software evolution. Rename refactoring is the most frequently performed refactoring that suggests a new name for an identifier to enhance readability when the identifier is poorly named. However, most existing works only identify renaming activities between two versions of source code, while few works express concern about how to suggest a new name. In this paper, we study automatic rename refactoring on variable names, which is considered more challenging than other rename refactoring activities. We first point out the connections between rename refactoring and various prevalent learning paradigms and the difference between rename refactoring and general text generation in natural language processing. Based on our observations, we propose RefBERT, a two-stage pre-trained framework for rename refactoring on variable names. RefBERT first predicts the number of sub-tokens in the new name and then generates sub-tokens accordingly. Several techniques, including constrained masked language modeling, contrastive learning, and the bag-of-tokens loss, are incorporated into RefBERT to tailor it for automatic rename refactoring on variable names. Through extensive experiments on our constructed refactoring datasets, we show that the generated variable names of RefBERT are more accurate and meaningful than those produced by the existing method. Our implementation and data are available at https://github.com/KDEGroup/RefBERT. Hao Liu 0003, Yanlin Wang 0001, Zhao Wei, Yong Xu 0007, Juhong Wang, Hui Li 0057, Rongrong Ji |
ISSTA | 3 |
| 2023 | EALink: An Efficient and Accurate Pre-Trained Framework for Issue-Commit Link RecoveryabstractIssue-commit links, as a type of software traceability links, play a vital role in various software development and maintenance tasks. However, they are typically deficient, as developers often forget or fail to create tags when making commits. Existing studies have deployed deep learning techniques, including pretrained models, to improve automatic issue-commit link recovery. Despite their promising performance, we argue that previous approaches have four main problems, hindering them from recovering links in large software projects. To overcome these problems, we propose an efficient and accurate pre-trained framework called EALink for issue-commit link recovery. EALink requires much fewer model parameters than existing pre-trained methods, bringing efficient training and recovery. Moreover, we design various techniques to improve the recovery accuracy of EALink. We construct a large-scale dataset and conduct extensive experiments to demonstrate the power of EALink. Results show that EALink outperforms the state-of-the-art methods by a large margin (15.23%-408.65%) on various evaluation metrics. Meanwhile, its training and inference overhead is orders of magnitude lower than existing methods. We provide our implementation and data at https://github.com/KDEGroup/EALink. Yanlin Wang 0001, Zhao Wei, Yong Xu 0007, Juhong Wang, Hui Li 0057, Rongrong Ji |
ASE | 3 |
| 2021 | Cross-language Code Coupling Detection: A Preliminary Study on Android ApplicationsabstractFramework-based multi-lingual software is increasingly prevalent, but it also brings negative effects and extra burden on software maintenance and evolution, because of the introduced cross-language code coupling, which are usually mixed with framework-specific conventions. Researchers have proposed various approaches to code coupling detection, but there is still a lack of necessary support for cross-language coupling detection in framework-based software development. In this paper, we present a preliminary study about cross-language coupling detection in software development based on the Android application framework. We investigate the characteristics of multi-lingual changes in the top-100 starred open-source Android repositories on GitHub, and find that multi-lingual commits are non-trivial: their code changes are more scattered, and more inclined to introduce bugs than other commits. To mitigate the side-effect of multi-lingual development, we propose Grace, a Graph-based cross-language co-change suggestion approach for Android application development. Grace (a) designs a language-agnostic graph to represent code elements from different languages, and (b) employs an entity-based collaborative filtering algorithm to detect and rank candidates of cross-language code couplings, from the graph representation of the latest version as well as the historical multi-lingual commits of a repository. To evaluate the effectiveness of Grace, we apply it to the two tasks of cross-language co-change suggestion and inconsistency checking. Results show that Grace (a) can effectively suggest cross-language co-changed files and types, and (b) can also find existing and potential bugs or code smells caused by inconsistent co-changes. Wei Zhang 0004, Ailun Yu, Zhao Wei, Guangtai Liang, Haiyan Zhao 0001, Zhi Jin 0001 |
ICSME | 4 |
| 2021 | SmartCommit: a graph-based interactive assistant for activity-oriented commitsabstractIn collaborative software development, it is considered to be a best practice to submit code changes as a sequence of cohesive commits, each of which records the work result of a specific development activity, such as adding a new feature, bug fixing, and refactoring. However, rather than following this best practice, developers often submit a set of loosely-related changes serving for different development activities as a composite commit, due to the tedious manual work and lack of effective tool support to decompose such a tangled changeset. Composite commits often obfuscate the change history of software artifacts and bring challenges to efficient collaboration among developers. To encourage activity-oriented commits, we propose SmartCommit, a graph-partitioning-based interactive approach to tangled changeset decomposition that leverages not only the efficiency of algorithms but also the knowledge of developers. To evaluate the effectiveness of our approach, we (1) deployed SmartCommit in an international IT company, and analyzed usage data collected from a field study with 83 engineers over 9 months; and (2) conducted a controlled experiment on 3,000 synthetic composite commits from 10 diverse open-source projects. Results show that SmartCommit achieves a median accuracy between 71–84% when decomposing composite commits without developer involvement, and significantly helps developers follow the best practice of submitting activity-oriented commits with acceptable interaction effort and time cost in real collaborative software development. Wei Zhang 0004, Christian Kästner, Haiyan Zhao 0001, Zhao Wei, Guangtai Liang, Zhi Jin 0001 |
ESEC/SIGSOFT FSE | 5 |
| 2019 | Dice Loss in Siamese Network for Visual Object Tracking
Zhao Wei, Changhao Zhang, Kaiming Gu, Fei Wang 0008 |
ICIC (2) | 1 |
| 2018 | Online Multi-Object Tracking with Structural Invariance Constraint
Peilin Jiang, Zhao Wei, Hang Dong 0001, Fei Wang 0008 |
BMVC | 3 |
| 2015 | Derivation of new split window algorithm for retrieving land surface temperature from FY-3/VIRR dataabstractLand surface temperature (LST) is a crucial parameter in analyzing and evaluating climate change at various scales, the surface energy balance, soil moisture, evapotranspiration and urban heat islands. Currently, methods for its estimation from space have continuously been developed, while most studies focus on the Split-Window (SW) algorithms. According to the published works, some approximations and assumptions were used to develop SW algorithms. This paper investigated and revised the error caused by these approximations and assumptions with the help of TIGR 2000 database and MODTRAN 4.0 software. Then a new SW method to estimate LST from FY-3A/VIRR was proposed in this paper. The primarily accuracy evaluation of the proposed method shows that the root mean square error (RMSE) of LST estimation using TIGR atmospheric profiles is 0.768 K, with the bias of −0.122 K. Sibo Duan, Zhao Wei, Zhao-Liang Li |
IGARSS | 3 |
| 2015 | Combination of contrast limited adaptive histogram equalisation and discrete wavelet transform for image enhancementabstractImage enhancement has an important role in image processing applications. Contrast limited adaptive histogram equalisation (CLAHE) is an effective algorithm to enhance the local details of an image. However, it faces the contrast overstretching and noise enhancement problems. To solve these problems, this study presents a novel image enhancement method, named CLAHE‐discrete wavelet transform (DWT), which combines the CLAHE with DWT. The new method includes three main steps: First, the original image is decomposed into low‐frequency and high‐frequency components by DWT. Then, the authors enhance the low‐frequency coefficients using CLAHE and keep the high‐frequency coefficients unchanged to limit noise enhancement. This is because the high‐frequency component corresponds to the detail information and contains most noises of original image. Finally, reconstruct the image by taking inverse DWT of the new coefficients. To alleviate over‐enhancement, the reconstructed and original images are averaged using an originally proposed weighting factor. The weighting operation can control the enhancement levels of regions with different luminances in original image adaptively. This is important because bright parts of image are usually needless to be enhanced in comparison with the dark parts. Extensive experiments show that this method performs well in detail preservation and noise suppression. Huang Lidong, Zhao Wei, Wang Jun, Zebin Sun |
IET Image Process. | 2 |
| 2015 | Modeling, conflict detection, and verification of a new virtualization role-based access control frameworkabstractAbstract In the last 10 years, virtualization has become a widespread technique in cloud computing; however, few of the access control models have ever addressed the security issue of multi‐domain and virtualized network management; this paper enhanced the classic role‐based access control model through two concepts: domain and virtual machine. We defined a new model named VRBAC in which authorized users can migrate or copy virtual machines from one domain to another without causing a conflict. Domain users or groups are allowed to share permissions of not only resources like shared files but also virtual machines with others either from the same or a different domain. Three kinds of VRBAC policy conflicts are defined in forms of ontologies, which provide extra access to description logic reasoning and facilitate the policy conflict detection. The experimental results based on Microsoft Active Directory and VMware vSphere suggest that all policy conflicts can be detected effectively and efficiently. Moreover, the generated reports can provide conflict details such as conflict types, positions, and causes, which will serve as guidance for further resolution of the improper authorizations and access violations. Copyright © 2014 John Wiley & Sons, Ltd. Chunhe Xia, Liangshuang Lv, Zhao Wei, Yazhuo Li |
Secur. Commun. Networks | 4 |
| 2012 | Chemical Reaction Optimization for the Fuzzy Rule learning problemabstractIn this paper, we utilize Chemical Reaction Optimization (CRO), a newly proposed metaheuristic for global optimization, to design Fuzzy Rule-Based Systems (FRBSs). CRO imitates the interactions of molecules in a chemical reaction. The molecular structure corresponds to a solution, and the potential energy is analogous to the objective function value. Molecules are driven toward the lowest energy stable state, which corresponds to the global optimum of the problem. In the realm of modeling with fuzzy rule-based systems, automatic derivation of fuzzy rules from numerical data plays a critical role. We propose to use CRO with Cooperative Rules (COR) to solve the fuzzy rule learning problem in FRBS. We formulate the learning process of FRBS in the form of a combinatorial optimization problem. Our proposed method COR-CRO is evaluated by two fuzzy modeling benchmarks and compared with other learning algorithms. Simulation results demonstrate that COR-CRO is highly competitive and outperforms many other existing optimization methods. Albert Y. S. Lam, Victor O. K. Li, Zhao Wei |
IEEE Congress on Evolutionary Computation | 3 |
| 2012 | Semantic Similarity Analysis between High-Level Model Description Text and Low-Level Implementation Text for Network SurvivabilityabstractThe transformation from high-level model description to low-level implementation for network survivability may lead to semantic inconsistency problem, since the method of transformation is based on symbolic transformation of machine which neglects semantics of proposed problem. Therefore, in order to analyze the semantic differences of transformation before and after and ensure the semantic consistency, this paper built an Ontology model of high-level model description and low-level implementation for network survivability. We proposed an Ontology-based multi-factor synthesized method which is used to analyze semantic similarity of multi-hierarchy text including high-level model description text and low-level implementation text. This method included the computation of type similarity of the text based on concept lattice, computation of content similarity of the text based on ontology and computation of structure similarity of the text based on inverse order pairs. The availability of the method is verified by experiments and its accuracy is proved by comparison with classic VSM algorithm, WeiSong algorithm, CF algorithm and professional value. Zhao Wei, Chunhe Xia, Shan Yao, Yujian Zhou |
Web Intelligence | 1 |