VLDB 2026 Research / reviewers in the wild / expert
Chun Yong Chong
dblp:134/9473
· DBLP profile ↗
32ranked-venue papers
5as first author
27since 2021 · last 2026
0000-0003-1164-0049ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 24 · 5 first-author · 19 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MulVul: Retrieval-augmented Multi-Agent Code Vulnerability Detection via Cross-Model Prompt EvolutionabstractLarge Language Models (LLMs) struggle to automate real-world vulnerability detection due to two key limitations: the heterogeneity of vulnerability patterns undermines the effectiveness of a single unified model, and manual prompt engineering for massive weakness categories is unscalable.To address these challenges, we propose MulVul, a retrievalaugmented multi-agent framework designed for precise and broad-coverage vulnerability detection.MulVul adopts a coarse-to-fine strategy: a Router agent first predicts the top-k coarse categories and then forwards the input to specialized Detector agents, which identify the exact vulnerability types.Both agents are equipped with retrieval tools to actively source evidence from vulnerability knowledge bases to mitigate hallucinations.Crucially, to automate the generation of specialized prompts, we design Cross-Model Prompt Evolution, a prompt optimization mechanism where a generator LLM iteratively refines candidate prompts while a distinct executor LLM validates their effectiveness.This decoupling mitigates the self-correction bias inherent in single-model optimization.Evaluated on 130 CWE types, MulVul achieves 34.79% Macro-F1, outperforming the best baseline by 41.5%.Ablation studies validate cross-model prompt evolution, which boosts performance by 51.6% over manual prompts by effectively handling diverse vulnerability patterns. Jie Xu 0031, Chun Yong Chong, Xiaohua Jia |
ACL (1) | 4 |
| 2026 | Beyond Coverage: Automatic Test Suite Augmentation for Enhanced Effectiveness using Large Language ModelsabstractLarge Language Models (LLMs) have gained significant traction in software engineering for automating tasks such as unit test generation. Most existing studies prioritize code coverage as the primary metric for enhancing test suite effectiveness. However, prior research has shown that although code coverage can reach approximately 80%, the mutation score, which generally exhibits a stronger correlation with defect detection effectiveness, attains only about 35%. This gap highlights the need to enhance test suite effectiveness guided by mutation score rather than code coverage. Recent studies, including MuTAP and Mu tGe n, explored the use of survived mutants to enhance test suite effectiveness. However, their evaluations were limited to simple standalone methods that rely on built-in functions and standard libraries. Non-standalone methods, which depend on other classes and involve complex user-defined types, are more intricate and commonly found in real-world projects. The limited contextual information and basic repair mechanisms in their prompt designs make it unclear whether their performance can generalize to non-standalone methods. Moreover, the two studies rely on existing language-specific, rule-based mutation techniques, which require specific configurations and incur additional costs when adapting to other programming languages. To bridge this gap, we propose a novel, fully automatic LLM-based approach to enhance test suite effec-tiveness, guided by survived mutants. The approach augments initial test suites by integrating mutation testing with test case generation. It takes focal method information as input and generates test cases targeting survived mutants identified from applying the initial test suites. Our approach incorporates multiple prompt techniques, rich contextual information, and an advanced repair mechanism to effectively generate test cases for non-standalone methods. The evaluation covers 1,035 focal methods, categorized as standalone or non-standalone. On average, the mutation score increases by 16.04% for standalone methods and 8.11% for non-standalone methods. We validate the practical impact of augmented test suites in LLM-based code generation. After test suite augmentation, pass@1 decreased by 0.3152 and 0.1772 on average for standalone and non-standalone methods, respectively, indicating the effectiveness of our approach in reducing false positives caused by insufficient test cases in code generation evaluation. Peng Zhang 0083, Yuge Nie, Yibiao Yang, Yutian Tang, Chun Yong Chong, Yuming Zhou |
Proc. ACM Program. Lang. | 6 |
| 2025 | Multi-view Leaderboard: Towards Evaluating the Code Intelligence of LLMs From Multiple ViewsabstractLarge Language Models (LLMs) have shown remarkable performance in code intelligence tasks, prompting the development of various benchmarks and leaderboards to assess their effectiveness across diverse programming scenarios. However, existing leaderboards often rely on coarse-grained metrics and overlook performance variations across different types of tasks. In this paper, we introduce Multi-view Leaderboard, a comprehensive evaluation framework designed to assess the coding capabilities of LLMs from multiple views. Our leaderboard partitions widely-used datasets such as HumanEval, MBPP, and ComplexCodeEval into subsets based on factors like prompt length, problem complexity, and task type. It supports four popular code intelligence tasks including code generation, code completion, test case generation, and API recommendation. Additionally, our leaderboard presents results using ranking tables, line charts, radar charts, and heatmaps. Based on LLMs’ performance on different subsets, we provide model recommendations tailored to different real-world scenarios via a Sankey diagram. A user study involving 11 participants revealed that 90% valued the leaderboard’s practical usefulness for analyzing LLMs’ code intelligence from multiple perspectives. The Multi-view Leaderboard is available at https://huggingface.co/spaces/MVLLL/Multi-view-leaderboard. The demonstration video is available at https://youtu.be/J-zQiOYa1Y8 Zexun Zhan, Cuiyun Gao 0001, Yujia Chen 0004, Guoai Xu, Chun Yong Chong, Shan Gao 0009, Xin Xia 0001 |
APSEC | 6 |
| 2025 | Diplomatist: What Do Cross-language Dependencies Reflect Software Ecosystem Health?abstractIn large-scale software development, multilingual projects, those involving multiple interacting programming languages, have become increasingly common in both industry and the open-source community. Research indicates that cross-language dependencies in these projects can increase the like-lihood of risks, such as functionality defects and security vulnerabilities. While most existing studies focus on cross-language dependencies between host languages and specific guest languages (e.g., C/C++), interactions between host languages and a broader range of guest languages, as well as the broader impact of such dependencies on software ecosystems, remain underexplored.To address the above limitations, in this paper, we develop a technique, Diplomatist, to identify and analyze cross-language dependencies between host languages, such as Java, and guest languages, including JavaScript, Python, Ruby, PHP, and C/C++. Diplomatist automatically analyzes cross-language invocation APIs and constructs a large-scale knowledge repository to standardize code features for identifying library versions across various guest languages, enabling host languages to trace the guest language libraries they invoke. Evaluation shows that Diplomatist achieved an average precision of 88.9% and a recall of 91.5% on a high-quality benchmark, indicating its high accuracy in detecting cross-language dependencies. Using Diplomatist, we identified 435,258 Java libraries that indirectly or transitively depend on libraries from other ecosystems. Diplomatist provides a list of cross-language pivotal libraries that contribute to preserving the long-term health and sustainability of software ecosystems. Moreover, we conduct a case study to examine the impact of the risks introduced due to cross-language dependencies on programming language ecosystems, by analyzing a full-picture of the cross-language dependency graph. Our findings show that fragile projects or libraries can propagate security issues across ecosystems via these dependencies, impacting 13,739 downstream projects in the Maven ecosystem. We utilized Diplomatist to provide remediation suggestions to relevant project developers. Issue reports of some subjects have been confirmed by developers. Fanyi Meng 0003, Ying Wang 0038, Chun Yong Chong, Hai Yu 0001, Zhiliang Zhu 0001 |
ASE | 3 |
| 2024 | LLM-Based Class Diagram Derivation from User Stories with Chain-of-Thought PromptingsabstractIn agile requirements engineering, user stories are the primary means of capturing project requirements. However, deriving conceptual models, such as class diagrams, from user stories requires significant manual effort. This paper explores the potential of leveraging Large Language Models (LLMs) and a tailored Chain-of- Thought (CoT) prompting technique to automate this task. We conducted a comprehensive preliminary study to investigate different prompting techniques applied to the task. The study involved comparing LLM-based approaches with guided and unguided human extraction to evaluate the effectiveness of the proposed LLM-based techniques. Our findings demonstrate that LLM-based approaches, particularly when combined with well-crafted few-shot prompts, outperform guided human extraction in identifying classes. However, we also identified areas of suboptimal performance through qualitative analysis. The proposed CoT prompting technique offers a promising pathway to automate the derivation of class diagrams in agile projects, reducing the reliance on manual effort. Our study contributes valuable insights and directions for future research in this field. Yishu Li, Jacky W. Keung, Chun Yong Chong, Yihan Liao |
COMPSAC | 4 |
| 2024 | CLUE: Customizing clustering techniques using machine learning for software modularizationabstractSoftware clustering is often used as a remodularization and architecture recovery technique to help developers simplify software maintenance tasks and ease the burden of software comprehension. While the choice of clustering technique can significantly influence the outcomes of remodularization, it is noteworthy that existing works have yet to conduct an exhaustive exploration of the suitability of various clustering techniques for different software projects. Although many prior works introduce new clustering techniques, their validations often focus on specific domains, which may restrict the generalizability of their findings. In this paper, we conduct an empirical study aimed at understanding the impact of software features and architectural problems on the effectiveness of various software clustering techniques. Leveraging our empirical findings, we propose an approach, CLUE, which leverages Machine Learning (ML) models to customize a suitable software clustering technique for a given software. Our approach focuses on eight types of software clustering techniques and offers insights into their suitability based on features and architectural problems of software. This comprehensive analysis helps developers to select the suitable clustering technique that can achieve the best MoJoFM, c2ccvg, or TurboMQ value from the chosen pool of software clustering techniques for specific software. We evaluate CLUE by analyzing 100 open-source software projects. The experiment results demonstrate that CLUE achieves highly accurate clustering technique customization, with an accuracy exceeding 90%. Fanyi Meng 0003, Ying Wang 0038, Chun Yong Chong, Hai Yu 0001, Zhiliang Zhu 0001 |
Internetware | 3 |
| 2024 | ComplexCodeEval: A Benchmark for Evaluating Large Code Models on More Complex CodeabstractIn recent years, with the widespread attention of academia and industry on the application of large language models (LLMs) to code-related tasks, an increasing number of large code models (LCMs) have been proposed and corresponding evaluation benchmarks have continually emerged. Although existing evaluation benchmarks are helpful for comparing different LCMs, they may not reflect the performance of LCMs in various development scenarios. Specifically, they might evaluate model performance in only one type of scenario (e.g., code generation or code completion), whereas real development contexts are diverse and may involve multiple tasks such as code generation, code completion, API recommendation, and test function generation. Additionally, the questions may not originate from actual development practices, failing to capture the programming challenges faced by developers during the development process. Cuiyun Gao 0001, Chun Yong Chong, Chaozheng Wang, Shan Gao 0009, Xin Xia 0001 |
ASE | 4 |
| 2024 | DataRecipe - How to Cook the Data for CodeLLM?abstractDespite the proliferation of language models, a lack of transparency persists regarding the training datasets used. Security concerns are often cited, but identifying high-quality training data is crucial for optimal model performance. Yet, while significant efforts have been made to improve model performance, dataset quality remains an under-explored area. Our study addresses this gap by comprehensively investigating data-quality properties and processing strategies used to train code generation models. We focus on identifying dataset features that impact model performance and leverage these insights to optimize datasets and enhance model efficacy. Our approach involves a multifaceted analysis encompassing metadata, statistics, data quality issues, semantic correlations between intent and code, and design choices. By manipulating these features, we explore their influence on model performance. Our findings reveal that dataset design choices significantly impact the performance of code generation models. Additionally, semantic correlations between intent and code can also affect performance, although to varying degrees. Kisub Kim, Jounghoon Kim, Byeongjo Park, Dongsun Kim 0001, Chun Yong Chong, Tiezhu Sun, Daniel Tang, Jacques Klein, Tegawendé F. Bissyandé |
ASE | 5 |
| 2024 | A Systematic Evaluation of Large Code Models in API Suggestion: When, Which, and HowabstractAPI suggestion is a critical task in modern software development, assisting programmers by predicting and recommending third-party APIs based on the current context. Recent advancements in large code models (LCMs) have shown promise in the API suggestion task. However, they mainly focus on suggesting which APIs to use, ignoring that programmers may demand more assistance while using APIs in practice including when to use the suggested APIs and how to use the APIs. To mitigate the gap, we conduct a systematic evaluation of LCMs for the API suggestion task in the paper. Chaozheng Wang, Shuzheng Gao, Cuiyun Gao 0001, Wenxuan Wang 0001, Chun Yong Chong, Shan Gao 0009, Michael R. Lyu |
ASE | 5 |
| 2024 | Siformer: Feature-isolated Transformer for Efficient Skeleton-based Sign Language RecognitionabstractSign language recognition (SLR) refers to interpreting sign language glosses from given videos automatically. This research area presents a complex challenge in computer vision because of the rapid and intricate movements inherent in sign languages, which encompass hand gestures, body postures, and even facial expressions. Recently, skeleton-based action recognition has attracted increasing attention due to its ability to handle variations in subjects and backgrounds independently. However, current skeleton-based SLR methods exhibit three limitations: 1) they often neglect the importance of realistic hand poses, where most studies train SLR models on non-realistic skeletal representations; 2) they tend to assume complete data availability in both training or inference phases, and capture intricate relationships among different body parts collectively; 3) these methods treat all sign glosses uniformly, failing to account for differences in complexity levels regarding skeletal representations. To enhance the realism of hand skeletal representations, we present a kinematic hand pose rectification method for enforcing constraints. Mitigating the impact of missing data, we propose a feature-isolated mechanism to focus on capturing local spatial-temporal context. This method captures the context concurrently and independently from individual features, thus enhancing the robustness of the SLR model. Additionally, to adapt to varying complexity levels of sign glosses, we develop an input-adaptive inference approach to optimise computational efficiency and accuracy. Experimental results demonstrate the effectiveness of our approach, as evidenced by achieving a new state-of-the-art (SOTA) performance on WLASL100 and LSA64. For WLASL100, we achieve a top-1 accuracy of 86.50%, marking a relative improvement of 2.39% over the previous SOTA. For LSA64, we achieve a top-1 accuracy of 99.84%. The artefacts and code related to this study are made publicly online (https://github.com/mpuu00001/Siformer.git). Muxin Pu, Mei Kuan Lim, Chun Yong Chong |
ACM Multimedia | 3 |
| 2024 | REARRANGE: Effort estimation approach for software clustering-based remodularisationabstractContext: Most research in software clustering and remodularisation typically concludes by recommending the refactoring operations without further insight into the practicality of the proposed technique. Developers might be hesitant to follow through with the refactoring suggestions due to the uncertainty in the effort needed. Objective: This work aims to address this gap by introducing an effoRt Estimation AppRoach foR softwAre clusteriNG-based rEmodularisation (REARRANGE) to close the loop in extant software clustering and remodularisation research by estimating the time required to carry out the suggested refactoring operations based on the history of the evolution of the software. By providing tangible estimates of refactoring effort in person-hours, we can inform developers of complex and time-consuming refactoring operations that will help prioritise refactoring efforts, allowing practitioners to weave in these activities during sprint planning. Method: REARRANGE builds a machine learning model to predict effort estimation based on past commit activity which extracts Software Features (lines of code, number of methods), Refactoring Features (refactoring type, source and destination) and Dependency Features (dependencies between classes). REARRANGE is then compared against sanity checks, baseline effort estimation models, and state-of-the-art software estimation models. We also attempt to cross-validate REARRANGE's effort estimation with software developers. Results: Experimented through 25 open-source Java-based projects, the proposed approach estimated the refactoring effort of the test subjects with a Mean Absolute Error (MAE) of 5.47 person-hours against the MAE of the next-best approach of 453.31 person-hours. Based on a survey conducted among software developers, REARRANGE consistently delivers accurate estimates in 93.6% of cases. Conclusion: The lack of a direct comparison for REARRANGE highlights the need for a refactoring effort-focused estimation model that provides tangible effort estimates in person-hours for refactoring operations. Only then can developers selectively choose relevant refactoring operations while considering the available time and budget constraints, bridging the gap between software clustering research and real-world application. Alvin Jian Jia Tan, Chun Yong Chong, Aldeida Aleti |
Inf. Softw. Technol. | 2 |
| 2024 | Automated construction of reference model for software remodularization through software evolutionabstractAbstract The undocumented evolution of a software project and its underlying architecture underscores the need to recover the architecture from the software's implementation‐level artifacts. Despite the existence of various software remodularization techniques, they often suffer from inaccuracies, and evaluating their effectiveness is challenging due to the absence of accurate “ground‐truth” architectures or reference models. Prior studies on reference model construction are time‐consuming and labor‐intensive as it heavily relies on manual analysis by domain experts. Besides, other existing approaches that directly utilize the directory or package structure of the latest version can be unreliable, lacking in‐depth analysis of the employed software structure. To address the above limitations, in this paper, we proposeAutomatedConstruction ofReferenceModel (ACRM), an approach for automatically constructing reference models by assigning weights to classes for various software projects using the metadata of all software versions and historical maintenance records. We evaluate ACRM through both quantitative and qualitative analyses. The experiment results provide quantitative validation and show that the generated reference models are reasonable, as confirmed by the relationship between proposed reference models and architectural smells or bugs. Furthermore, we conduct a survey among the practitioners from industry, to gain insights from practitioners' practices and further validate the generated reference models. The survey shows that, on average, 87% of the participants agree with the reference models generated by ACRM. Moreover, we propose an improved metric,wc2c, which analyzes the strengths and weaknesses of different types of software clustering techniques using the proposed reference models of the given software. Finally, we discuss the potential benefits of using ACRM in analyzed projects, particularly in terms of improving software quality, reducing maintenance costs, and enhancing developer productivity. Fanyi Meng 0003, Hai Yu 0001, Chun Yong Chong, Ying Wang 0038, Zhiliang Zhu 0001 |
J. Softw. Evol. Process. | 3 |
| 2024 | Evolution-Aware Constraint Derivation Approach for Software RemodularizationabstractExisting software clustering techniques tend to ignore prior knowledge from domain experts, leading to results (suggested big-bang remodularization actions) that cannot be acceptable to developers. Incorporating domain experts knowledge or constraints during clustering ensures the obtained modularization aligns with developers’ perspectives, enhancing software quality. However, manual review by knowledgeable domain experts for constraint generation is time-consuming and labor-intensive. In this article, we propose an evolution-aware constraint derivation approach, Escort , which automatically derives clustering constraints based on the evolutionary history from the analyzed software. Specifically, Escort can serve as an alternative approach to derive implicit and explicit constraints in situations where domain experts are absent. In the subsequent constrained clustering process, Escort can be considered as a framework to help supplement and enhance various unconstrained clustering techniques to improve their accuracy and reliability. We evaluate Escort based on both quantitative and qualitative analysis. In quantitative validation, Escort , using generated clustering constraints, outperforms seven classic unconstrained clustering techniques. Qualitatively, a survey with developers from five IT companies indicates that 89% agree with Escort ’s clustering constraints. We also evaluate the utility of refactoring suggestions from our constrained clustering approach, with 54% acknowledged by project developers, either implemented or planned for future releases. Fanyi Meng 0003, Ying Wang 0038, Chun Yong Chong, Hai Yu 0001, Zhiliang Zhu 0001 |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2023 | ASDF: A Differential Testing Framework for Automatic Speech Recognition SystemsabstractRecent years have witnessed wider adoption of Automated Speech Recognition (ASR) techniques in various domains. Consequently, evaluating and enhancing the quality of ASR systems is of great importance. This paper proposes Asdf, an Automated Speech Recognition Differential Testing Framework to test ASR systems. Asdf extends an existing ASR testing tool, the CrossASR++, which synthesizes test cases from a text corpus. However, CrossASR++ fails to make use of the text corpus efficiently and provides limited information on how the failed test cases can improve ASR systems. To address these limitations, our tool incorporates two novel features: (1) a text transformation module to boost the number of generated test cases and uncover more errors in ASR systems, and (2) a phonetic analysis module to identify phonemes that the ASR systems tend to transcribe incorrectly. Asdf generates more high-quality test cases by applying various text transformation methods (e.g., changing tense) to the input text in a failed test case. By doing so, Asdf can utilize a small text corpus to generate a large number of audio test cases, something which CrossASR++ is not capable of. In addition, Asdf implements more metrics to evaluate the performance of ASR systems from multiple perspectives. Asdf performs phonetic analysis on the identified failed test cases to identify the phonemes that ASR systems tend to transcribe incorrectly, providing useful information for developers to improve ASR systems. The demonstration video of our tool is made online at https://www.youtube.com/watch?v=DzVwfc3h9As. The implementation is available at https://github.com/danielyuenhx/asdf-differential-testing. Daniel Hao Xian Yuen, Andrew Yong Chen Pang, Zhou Yang 0003, Chun Yong Chong, Mei Kuan Lim, David Lo 0001 |
ICST | 4 |
| 2023 | Synthesizing Speech Test Cases with Text-to-Speech? An Empirical Study on the False Alarms in Automated Speech Recognition TestingabstractRecent studies have proposed the use of Text-To-Speech (TTS) systems to automatically synthesise speech test cases on a scale and uncover a large number of failures in ASR systems. However, the failures uncovered by synthetic test cases may not reflect the actual performance of an ASR system when it transcribes human audio, which we refer to as false alarms. Given a failed test case synthesised from TTS systems, which consists of TTS-generated audio and the corresponding ground truth text, we feed the human audio stating the same text to an ASR system. If human audio can be correctly transcribed, an instance of a false alarm is detected. Julia Kaiwen Lau, Kelvin Kai Wen Kong, Julian Hao Yong, Per Hoong Tan, Zhou Yang 0003, Zi Qian Yong, Joshua Chern Wey Low, Chun Yong Chong, Mei Kuan Lim, David Lo 0001 |
ISSTA | 8 |
| 2023 | Exploring and Repairing Gender Fairness Violations in Word Embedding-based Sentiment Analysis Model through Adversarial PatchesabstractWith the advancement of sentiment analysis (SA) models and their incorporation into our daily lives, fairness testing on these models is crucial, since unfair decisions can cause discrimination to a large population. Nevertheless, some challenges in fairness testing include the unknown oracle, the difficulty in generating suitable test inputs, and the lack of a reliable way of fixing the issues. To fill in these gaps, BiasRV, a tool based on metamorphic testing (MT), was introduced and succeeded in uncovering fairness issues in a transformer-based model. However, the extent of unfairness in other SA models has not been thoroughly investigated. Our work conducts a more comprehensive empirical study to reveal the extent of fairness violations, specifically gender fairness, exhibited by other popular word embedding-based SA models. We define fairness violation as the behavior in which an SA model predicts variants created from a text, which merely differ in gender classes, to have different sentiments. Our inspection utilizing BiasRV uncovers at least 30 fairness violations (at BiasRV’s default threshold) in all three SA models. Realizing the importance of addressing such significant violations, we introduce adversarial patches (AP) as a way of patch generation in an automated program repair (APR) system to fix them. We adopt adversarial fine-tuning in AP by retraining SA models using adversarial examples, which are bias-uncovering test cases dynamically generated by a tool named BiasFinder at runtime. Evaluation of the SA models shows that our proposed AP reduces fairness violations by at least 25%. Lin Sze Khoo, Jia Qi Bay, Ming Lee Kimberly Yap, Mei Kuan Lim, Chun Yong Chong, Zhou Yang 0003, David Lo 0001 |
SANER | 5 |
| 2023 | Extended Abstract of E-SC4R: Explaining Software Clustering for RemodularisationabstractMaintenance of existing software requires a large amount of time for comprehending the source code. The architecture of a software, however, may not be clear to maintainers if up-to-date documentations are not available. Software clustering is often used as a remodularisation and architecture recovery technique to help recover a semantic representation of the software design. However, due to the diverse domain and structure of software systems, the suitability of different clustering techniques for different software systems are not investigated thoroughly. Research that introduce new clustering techniques usually validate their approaches on a specific domain, which might limit its generalisability. If the chosen test subjects only represent a narrow perspective of the whole picture, researchers risk not being able to address the external validity of their findings. This work aims to fill this gap by introducing a new approach, Explaining Software Clustering for Remodularisation (E-SC4R), to evaluate the effectiveness of different software clustering approaches. This work focuses on hierarchical clustering and Bunch clustering algorithms and provides information about their suitability according to the features of the software, which, as a consequence, enables the selection of the optimum technique for a particular software system. The E-SC4R framework is able to characterise both the strengths and weaknesses of the analysed software clustering algorithms using software features extracted from the code. The proposed approach also provides a better understanding of the algorithms’ behaviour by showing a 2D representation of the effectiveness of clustering techniques on the feature space through the application of dimensionality reduction techniques. Alvin Jian Jia Tan, Chun Yong Chong, Aldeida Aleti |
SANER | 2 |
| 2023 | General Outage Probability Model for UAV-to-UAV links in Multi-UAV Networks
Wei Jian Lau, Joanne Mun-Yee Lim, Chun Yong Chong, Nee Shen Ho, Thomas Wei Min Ooi |
Comput. Networks | 3 |
| 2023 | CodeLabeller: A Web-Based Code Annotation Tool for Java Design Patterns and SummariesabstractWhile constructing supervised learning models, we require labeled examples to build a corpus and train a machine learning model. However, most studies have built the labeled dataset manually, which, on many occasions, is a daunting task. To mitigate this problem, we have built an online tool called CodeLabeller. CodeLabeller is a web-based tool that aims to provide an efficient approach to handling the process of labeling source code files for supervised learning methods at scale by improving the data collection process throughout. CodeLabeller is tested by constructing a corpus of over a thousand source files obtained from a large collection of open source Java projects and labeling each Java source file with their respective design patterns and summaries. Twenty-five experts in the field of software engineering participated in a usability evaluation of the tool using the standard User Experience Questionnaire online survey. The survey results demonstrate that the tool achieves the Good standard on hedonic and pragmatic quality standards, is easy to use and meets the needs of annotating the corpus for supervised classifiers. Apart from assisting researchers in crowdsourcing a labeled dataset, the tool has practical applicability in software engineering education and assists in building expert ratings for software artefacts. Najam Nazar, Norman Chen, Chun Yong Chong |
Int. J. Softw. Eng. Knowl. Eng. | 3 |
| 2023 | A metadata driven process for assessing stability and reusability based on risk of change of software systemsabstractAbstract Measuring and estimating the reusability of software components are important steps toward finding reusable candidates. Reuse of software components can aid in the reduction of the development cost of new software or maintenance of existing ones. However, assessing the reusability of software is a challenging task. Even if reusable candidates can be identified, developers will need to decide on the reuse strategies, that is, reuse as it is without any modification, or introduce new changes to the reusable candidate to fulfill requirements in the new environment or releases. Thus, the stability of the reusable candidates plays a pivotal role because it will influence the ease of reuse. Assessing and predicting the stability of potential reusable candidates becomes a significant and important step in software reuse. In this research, we propose to leverage on risk of change as a proxy to measure software stability, where the latter is also a proxy measure for software reuse. Based on experiments conducted on 29 open‐source software with different sizes and application types hosted on GitHub, we found that classes with high impact and high risk of change should be avoided from being reused due to their instability, while those classes with low impact and risk of change should be given priority. The proposed work can aid in providing a better understanding of the ease of reuse for software systems and can be used as a tool to assess its overall quality. Bilal Mehboob, Chun Yong Chong |
Softw. Pract. Exp. | 2 |
| 2023 | High Dimensional Origin Destination Calibration Using Metamodel Assisted Simultaneous Perturbation Stochastic ApproximationabstractThe huge traffic data generated by intelligent transportation system (ITS) leads to the development of many advanced traffic models. These traffic models consist of many adjustable parameters which need to be calibrated before they are used in practice. This paper focuses on the offline calibration of origin-destination (OD) input parameters of a simulation-based traffic model to match its output with sensor data. An improved Simultaneous Perturbation Stochastic Approximation (SPSA) algorithm called Metamodel Assisted Simultaneous Perturbation Stochastic Approximation (MSPSA) is proposed in this paper to calibrate high-dimensional OD parameters within a tight computational budget. The proposed MSPSA combines the gradient of SPSA with the gradient of a differentiable metamodel function to improve the calibration efficiency. An integer program is also used to fine-tune the OD estimates. The proposed MSPSA algorithm is tested on a simple synthetic toy network and complex road network of Kuala Lumpur (KL), Malaysia. The proposed MSPSA algorithm is compared against SPSA and state-of-the-art Weighted Simultaneous Perturbation Stochastic Approximation (WSPSA) in both transportation networks. For KL network, synthetic and real-world sensor measurements are used as ground truth references to evaluate the performance of each approach. Based on the simulation results, the proposed MSPSA algorithm is able to gain at least 50% of improvement as compared to SPSA and WSPSA in both synthetic and real-world scenarios. Mun Chon Ho, Joanne Mun-Yee Lim, Chun Yong Chong, Kah Keong Chua, Alvin Kuok Lim Siah |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2023 | Collaborative Vehicle Rerouting System With Dynamic Vehicle SelectionabstractTraffic congestion represents a prevailing challenge encountered in urban areas due to the substantial volume of vehicles. To alleviate traffic congestion during rush hours, an improved vehicle rerouting strategy designed for vehicular ad hoc networks (VANETs) is proposed in this paper. Prior studies have predominantly relied on the number of hops upstream from a congested road to select vehicles to be rerouted away from traffic congestion. Selecting vehicles in this way causes the distance between the selected vehicles and the congested road to become sensitive to the network topology, especially when the road lengths vary significantly. Different from the previous research, the proposed vehicle rerouting strategy considers travel time, which is a dynamic traffic information, in selecting vehicles to be rerouted. Furthermore, to achieve collaborative rerouting among vehicles, the proposed vehicle rerouting strategy updates the remaining capacity of roads each time a new route is assigned to a selected vehicle. This is to ensure that the roads are not overutilized and become congested in the near future. The performance of the proposed rerouting strategy is compared with other state-of-the-art strategies in terms of average travel time, average carbon dioxide (CO2) emission, and average fuel consumption through traffic simulations in two transportation networks, which are simple grid network and real-world Kuala Lumpur (KL) network. Based on the simulation results, the proposed vehicle rerouting strategy outperforms other strategies by at least 4.39% in Kuala Lumpur (KL) network in terms of average travel time. Mun Chon Ho, Joanne Mun-Yee Lim, Chun Yong Chong, Kah Keong Chua, Alvin Kuok Lim Siah |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2022 | Metamorphic Testing-based Adversarial Attack to Fool Deepfake DetectorsabstractDeepfakes utilise Artificial Intelligence (AI) techniques to create synthetic media where the likeness of one person is replaced with another. There are growing concerns that deepfakes can be maliciously used to create misleading and harmful digital contents. As deepfakes become more common, there is a dire need for deepfake detection technology to help spot deepfake media. Present deepfake detection models are able to achieve outstanding accuracy (>90%). However, most of them are limited to within-dataset scenario. Most models do not generalise well enough in cross-dataset scenario. Furthermore, state-of-the-art deepfake detection models rely on neural network-based classification models that are known to be vulnerable to adversarial attacks. Motivated by the need for a robust deepfake detection model, this study adapts metamorphic testing (MT) principles to help identify potential factors that could influence the robustness of the examined model, while overcoming the test oracle problem in this domain. Metamorphic testing is specifically chosen as the testing technique as it fits our demand to address learning-based system testing with probabilistic outcomes from largely black-box components, based on potentially large input domains. We performed our evaluations on MesoInception-4 and TwoStreamNet models, which are the state-of-the-art deepfake detection models. This study identified makeup application as an adversarial attack that could fool deepfake detectors. Our experimental results demonstrate that both the MesoInception-4 and TwoStreamNet models degrade in their performance by up to 30% when the input data is perturbed with makeup. Nyee Thoang Lim, Meng Yi Kuan, Muxin Pu, Mei Kuan Lim, Chun Yong Chong |
ICPR | 5 |
| 2022 | A framework of dynamic selection method for user classification in touch-based continuous mobile device authentication
Ahmad Zairi bin Zaidi, Chun Yong Chong, Rajendran Parthiban, Ali Safa Sadiq |
J. Inf. Secur. Appl. | 2 |
| 2022 | E-SC4R: Explaining Software Clustering for Remodularisation
Alvin Jian Jia Tan, Chun Yong Chong, Aldeida Aleti |
J. Syst. Softw. | 2 |
| 2021 | Touch-based continuous mobile device authentication: State-of-the-art, challenges and opportunities
Ahmad Zairi bin Zaidi, Chun Yong Chong, Zhe Jin 0001, Rajendran Parthiban, Ali Safa Sadiq |
J. Netw. Comput. Appl. | 2 |
| 2021 | Reusability affecting factors and software metrics for reusability: A systematic literature reviewabstractAbstract Measuring and estimating the reusability of software components is important towards finding reusable candidates. Researchers have shown that software metrics can be effectively used to assess software reusability. This work provides a systematic literature review to investigate the main factors that influence software reusability and how these identified factors can be quantified using software metrics. This paper also investigates tool availability of the identified software metrics. Based on the extensive study, we narrowed down 44 factors that could positively or negatively affect the reusability of software systems. In term of software metrics, we report our findings through five main families of metrics, namely coupling, cohesion, complexity, inheritance, and size. We found that most of the metrics examine reusability at the class‐level, and the availability of software tools is limited. Furthermore, not all reusability affecting factors are equally impactful to assess the reusability of software components. While existing studies often discussed the impact of complexity towards software reusability, we found that only a handful of complexity metrics were meant for assessing reusability. We have identified several open challenges and gaps in the area, in particular lack of quantifiable measurement for reusability, limited software tools, and limited metrics that directly measure reusability. Bilal Mehboob, Chun Yong Chong, Sai Peck Lee, Joanne Mun-Yee Lim |
Softw. Pract. Exp. | 2 |
| 2018 | A Commit Change-based Weighted Complex Network Approach to Identify Potential Fault Prone Classes
Chun Yong Chong, Sai Peck Lee |
ICSOFT | 1 |
| 2017 | Automatic clustering constraints derivation from object-oriented software using weighted complex network with graph theory analysis
Chun Yong Chong, Sai Peck Lee |
J. Syst. Softw. | 1 |
| 2015 | Constrained Agglomerative Hierarchical Software Clustering with Hard and Soft ConstraintsabstractAlthough agglomerative hierarchical software clustering technique has been widely used in reverse engineering to recover a high-level abstraction of the software in the case of limited resources, there is a lack of work in this research context to integrate the concept of pair-wise constraints, such as must-link and cannot-link constraints, to further improve the quality of clustering. Pair-wise constraints that are derived from experts or software developers, provide a means to indicate whether a pair of software components belongs to the same functional group. In this paper, a constrained agglomerative hierarchical clustering algorithm is proposed to maximize the fulfilment of must-link and cannot-link constraints in a unique manner. Two experiments using real-world software systems are performed to evaluate the effectiveness of the proposed algorithm. The result of evaluation shows that the proposed algorithm is capable of handling constraints to improve the quality of clustering, and ultimately provide a better understanding of the analyzed software system. Chun Yong Chong, Sai Peck Lee |
ENASE | 1 |
| 2015 | Analyzing maintainability and reliability of object-oriented software using weighted complex network
Chun Yong Chong, Sai Peck Lee |
J. Syst. Softw. | 1 |
| 2013 | Efficient software clustering technique using an adaptive and preventive dendrogram cutting approach
Chun Yong Chong, Sai Peck Lee, Teck Chaw Ling |
Inf. Softw. Technol. | 1 |