EDBT 2026 Demo / reviewers in the wild / expert
Khanh Hoa Dam
dblp:75/4194 · also Hoa Khanh Dam
· DBLP profile ↗
74ranked-venue papers
17as first author
29since 2021 · last 2025
0000-0003-4246-0526ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 50 · 11 first-author · 22 since 2021Artificial intelligence and machine learning · 17 · 5 first-author · 6 since 2021Databases, data management, data science and information retrieval · 13 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Coordinated Self-Exploration for Self-Adaptive Systems in Contested Environments
Saad Sajid Hashmi, Khanh Hoa Dam, Alan W. Colman, Anton V. Uzunov, Quoc Bao Vo, Mohan Baruwal Chhetri, James Dorevski |
ICAART (1) | 2 |
| 2025 | An LLM-based multi-agent framework for agile effort estimationabstractEffort estimation is a crucial activity in agile software development, where teams collaboratively review, discuss, and estimate the effort required to complete user stories in a product backlog. Current practices in agile effort estimation heavily rely on subjective assessments, leading to inaccuracies and inconsistencies in the estimates. While recent machine learning-based methods show promising accuracy, they cannot explain or justify their estimates and lack the capability to interact with human team members. Our paper fills this significant gap by leveraging the powerful capabilities of Large Language Models (LLMs). We propose a novel LLM-based multi-agent framework for agile estimation that not only can produce estimates, but also can coordinate, communicate and discuss with human developers and other agents to reach a consensus. Evaluation results on a real-life dataset show that our approach outperforms state-of-the-art techniques across all evaluation metrics in the majority of cases. Our human study with software development practitioners also demonstrates an overwhelmingly positive experience in collaborating with our agents in agile effort estimation. Thanh-Long Bui, Khanh Hoa Dam, Rashina Hoda |
ASE | 2 |
| 2025 | Venturing ChatGPT's lens to explore human values in software artifacts: a case study of mobile APIsabstractSoftware is designed for humans and must account for their values. However, current research and practice focus on a narrow range of well-explored values, e.g. security, overlooking a more comprehensive perspective. Those exploring a broader array of values rely on manual identification, which is labour-intensive and prone to human bias. Moreover, existing methods offer limited reliability as they fail to explain their findings. In this paper, we propose leveraging the reasoning capabilities of Large Language Models (LLMs) for automated inference about values. This allows for not only detecting values but also explaining how they are expressed in the software. We aim to examine the effectiveness of LLMs, specifically ChatGPT (Chat Generative Pre-Trained Transformer), in automated detection and explanation of values in software artifacts. Using ChatGPT, we investigate how mobile APIs align with human values based on their documentation. Human evaluation of ChatGPT's findings shows a reciprocal shift in understanding values, with both ChatGPT and experts adjusting their assessments through dialogue. While experts recognise ChatGPT's potential for revealing values, emphasis is placed on human involvement to enhance the accuracy of the findings by detecting and eliminating convincing but inaccurate explanations provided by the language model due to potential hallucinations or confabulations. Davoud Mougouei, Saima Rafi, Mahdi Fahmideh, Elahe Mougouei, Javed Ali Khan, Khanh Hoa Dam, Arif Nurwidyantoro, Michel R. V. Chaudron |
Behav. Inf. Technol. | 6 |
| 2025 | Proactive self-exploration: Leveraging information sharing and predictive modelling for anticipating and countering adversaries
Saad Sajid Hashmi, Khanh Hoa Dam, Mohan Baruwal Chhetri, Anton V. Uzunov, Alan W. Colman, Quoc Bao Vo |
Expert Syst. Appl. | 2 |
| 2025 | Graph-based explainable vulnerability prediction
Hong Quy Nguyen, Thong Hoang, Khanh Hoa Dam, Aditya Ghose |
Inf. Softw. Technol. | 3 |
| 2025 | Human-understandable explanation for software vulnerability predictionabstractRecent advances in deep learning have significantly improved the performance of software vulnerability prediction (SVP). To enhance trustworthiness, the SVP highlights predicted lines of code (LoC) that may be vulnerable. However, providing LoC alone is often insufficient for software practitioners, as it lacks detailed information about the nature of the vulnerability. This paper introduces a novel framework that is built on SVP by offering additional explanatory information based on the suggested LoC. Similar to security reports, our framework comprehensively explains the vulnerability aspects, such as Root Cause, Impact, Attack Vector, and Vulnerability Type. The proposed framework is powered by transformer architectures. Specifically, we leverage pre-trained language models for code to fine-tune on two practical datasets: BigVul and Vulnerability Key Aspect, ensuring our framework’s applicability to real-world scenarios. Experiments using the ROUGE and BLEU scores as evaluation metrics show that our framework achieves better performance with CodeT5+, statistically outperforming a baseline study in generating key vulnerability aspects. Additionally, we conducted a small-scale user study with experienced software practitioners to assess the effectiveness of the framework. The results show that 72% of the participants found our framework helpful in accepting the SVP results, and 68% rated the additional explanations as moderately to extremely useful. Editor’s note: Open Science material was validated by the Journal of Systems and Software Open Science Board . • A novel framework generates an explanation from the predicted vulnerable lines. • Comprehensive investigations of factors influencing the quality of the framework. • We conducted a user study to validate its usefulness. Hong Quy Nguyen, Thong Hoang, Khanh Hoa Dam, Guoxin Su, Zhenchang Xing, Qinghua Lu 0001, Jiamou Sun |
J. Syst. Softw. | 3 |
| 2025 | Adversarial patch generation for automated program repair
Abdulaziz Alhefdhi, Khanh Hoa Dam, Thanh Le-Cong, Bach Le 0001, Aditya Ghose |
Softw. Qual. J. | 2 |
| 2025 | Sprint2Vec: A Deep Characterization of Sprints in Iterative Software DevelopmentabstractIterative approaches like Agile Scrum are commonly adopted to enhance the software development process. However, challenges such as schedule and budget overruns still persist in many software projects. Several approaches employ machine learning techniques, particularly classification, to facilitate decision-making in iterative software development. Existing approaches often concentrate on characterizing a sprint to predict solely productivity. We introduce Sprint2Vec, which leverages three aspects of sprint information – sprint attributes, issue attributes, and the developers involved in a sprint, to comprehensively characterize it for predicting both productivity and quality outcomes of the sprints. Our approach combines traditional feature extraction techniques with automated deep learning-based unsupervised feature learning techniques. We utilize methods like Long Short-Term Memory (LSTM) to enhance our feature learning process. This enables us to learn features from unstructured data, such as textual descriptions of issues and sequences of developer activities. We conducted an evaluation of our approach on two regression tasks: predicting the deliverability (i.e., the amount of work delivered from a sprint) and quality of a sprint (i.e., the amount of delivered work that requires rework). The evaluation results on five well-known open-source projects (Apache, Atlassian, Jenkins, Spring, and Talendforge) demonstrate our approach's superior performance compared to baseline and alternative approaches. Morakot Choetkiertikul, Peerachai Banyongrakkul, Chaiyong Ragkhitwetsagul, Suppawong Tuarob, Khanh Hoa Dam, Thanwadee Sunetnanta |
IEEE Trans. Software Eng. | 5 |
| 2024 | Microcompositions for Goal-Driven Self-AdaptationabstractTo adapt to volatile edge environments this paper envisages distributed systems built from small units of coordi-nation called microcompositions. Microcompositions decompose a single goal into subgoals and declaratively express compo-sition and coordination constraints. These microcompositions are managed by agents that self-organise themselves into a distributed management overlay structure - a Goal Realisation Tree (GRT) - based on goal realisation (means-ends) links. Both microcompositions and GRTs serve as runtime models, allowing agents to dynamically restructure them in response to changes in goals, service provision and context. Results from our evaluation demonstrate the effectiveness and scalability of our approach. Alan W. Colman, Quoc Bao Vo, Anton V. Uzunov, Saad Sajid Hashmi, Khanh Hoa Dam, Mohan Baruwal Chhetri |
COMPSAC | 5 |
| 2024 | Verifying Multi -Agent Coordination Correctness for BDI AgentsabstractThe popular Belief-Desire-Intention (BDI) architecture structures agents' behavior around beliefs, desires, and intentions, enabling effective pursuit of collective objectives and optimal system performance. Many complex situations require agent developers to design a coordination model/plan involving multiple BDI agents working together to achieve a given goal. However, a semantic notion of state has been missing in most current accounts of BDI agent execution, making it impossible to assess at design time if a coordination model/plan is able to achieve the goal. Thus, in this paper, we introduce a framework that permits us to identify the state that accrues at any point in agent execution. This enables us to determine, at design time, if a multi-agent coordination model/plan accomplishes a given goal. Our framework facilitates the verification of coordination correctness in such complex multi-agent systems. Geeta Mahala, Aditya Ghose, Khanh Hoa Dam, Angela Consoli |
COMPSAC | 3 |
| 2024 | MicroRec: Leveraging Large Language Models for Microservice RecommendationabstractThe increasing adoption of microservices in software development requires effective recommendation systems that guide developers to relevant microservices. In this paper, we introduce MicroRec, a novel microservice recommender framework which leverages insights from Stack Overflow posts and the power of Large Language Models (LLMs). MicroRec utilizes a dual-encoder architecture that combines contrastive learning and semantic similarity learning, allowing us to achieve robust and accurate retrieval and ranking of relevant posts based on user queries. Using LLMs, MicroRec builds up a deep understanding of both user queries and microservices through the information they provide (e.g., README files and Dockerfiles). Our empirical evaluations demonstrate significant improvements brought by MicroRec over the existing methods across a variety of performance metrics including MRR, MAP, and precision@k. In addition, the results returned by MicroRec were fourteen times more accurate than those provided by the existing recommendation tool on the widely-used Docker Hub platform. Ahmed Saeed Alsayed, Khanh Hoa Dam |
MSR | 2 |
| 2024 | Towards automating self-admitted technical debt repaymentabstractSelf-Admitted Technical Debt (SATD) refers to the technical debt in software that is explicitly flagged, typically by the source code comment. The SATD literature has mainly focused on comprehending, describing, detecting, and recommending SATD. Most recently, there have been efforts to study the state of the code before and after removing the SATD comment. While these efforts serve as a preliminary step towards the repayment of SATD, actual attempts towards automating SATD repayment, to the best of our knowledge, are yet to be made. In this paper, we propose the first attempt towards direct, complete, and automated SATD repayment by providing two main contributions. The first contribution is an empirical study of how the SATD comment relates to repaying the debt. The second contribution is DLRepay, our deep learning approach for SATD repayment. We developed a SATD Repayment dataset, namely SATD-R, and established a taxonomy based on the relationship and helpfulness of the SATD comment to/in repaying the debt. In addition, we developed DLRepay which takes as an input a pair of SATD comment and code, and generates a new, TD-free code. We found that there are five different categories in which the SATD comment relates to Technical Debt repayment. We also identify when the SATD comment has a positive and logical connection to repaying the debt, both generally and in every category. Furthermore, we illustrate the results of our SATD repayment approach across two datasets, three input types, two output types, and two neural networks. The resulting taxonomy of our empirical study paves the way for research to tackle further in-depth questions concerning SATD repayment comprehension, identification, and automation. In addition, the various experimental setups we conduct provide multiple insights regarding the applicability of our SATD repayment approach. Abdulaziz Alhefdhi, Khanh Hoa Dam, Aditya Ghose |
Inf. Softw. Technol. | 2 |
| 2024 | StraAlgin: Automated Strategic Alignment of ServicesabstractAligning strategic objectives with business services is important for an organisation's long-term success. The automation of those alignments offers significant benefits in terms of identifying strategies that are not supported by any services, or services that are irrelevant to the organisation's strategies. Automated strategy-service alignment also enables organizations to proactively respond to changes in complex environments and allocate resources effectively to increase the likelihood of achieving strategic objectives. StraAlgin offers a methodology and tool for defining, modeling and refining business objectives and strategies, and aligning them with business services. It also enables automated identification of the optimal choice of services that could realize organizational strategies. Evaluation results have demonstrated a high level accuracy of our approach and its scalability to a large number of strategies and services. This paper proposes an approach to automate the assessment and alignment of services and strategies. Our approach is implemented in the form of StraAlign toolkit. StraAlign provides the necessary vocabulary for business objective description and offers automated support to improve the evaluation and implementation of strategic service alignments. Ahmed Saeed Alsayed, Khanh Hoa Dam, Aditya Ghose |
IEEE Trans. Serv. Comput. | 2 |
| 2024 | Multi-Objective Evolutionary Search for Optimal Robotic Process Automation ArchitecturesabstractRobotic Process Automation (RPA) design and implementation requires an architecture which facilitates the seamless transition between human agents, robotic agents, and intelligent agents to automate information acquisition tasks and decision-making tasks. Coordination of those agents must consider various factors, such as efficiency of a resource when completing tasks, the quality of completed complex tasks, and the cost of the used resources. This article proposes a novel approach for generating an optimal architecture based on distinct types of resources, including human agents, intelligent agents, and robotic agents. An optimal architecture is the optimal enactment of process instances executed by a combination of human and automation agents based on their characteristics. The architecture provides a set of resources and their characteristics that are tailored to meet multiple objectives for process execution. The proposed approach is validated through an empirical evaluation based on a real-world business process. An empirical evaluation demonstrates that, given equal computational time, our approach outperforms conventional constraint optimization ILOG CPLEX (Manual 1987). Geeta Mahala, Renuka Sindhgatta, Khanh Hoa Dam, Aditya Ghose |
IEEE Trans. Serv. Comput. | 3 |
| 2023 | Goal-Driven Adversarial Search for Distributed Self-Adaptive SystemsabstractResilience and antifragility are highly desirable properties for systems operating in dynamic, contested environments. Recent work has proposed an approach for achieving antifragility through three new self-* properties, one of which is adversarial self-exploration. In that approach, a single agent is responsible for defending a (distributed) managed system against an adversary. However, a major limitation is that the agent requires full observability of the managed system and its environment, which is not practical in many scenarios - including when a system operates in contested environments. We address this limitation by extending the approach to multi-agent self-exploration, where agents with partial observability and control of the managed system coordinate with other agents to compute the most resilient responses in the presence of adversaries. We demonstrate the feasibility and scalability of the proposed approach through extensive experimentation. Saad Sajid Hashmi, Khanh Hoa Dam, Anton V. Uzunov, Mohan Baruwal Chhetri, Aditya Ghose, Alan W. Colman |
SSE | 2 |
| 2023 | On Privacy Weaknesses and Vulnerabilities in Software SystemsabstractIn this digital era, our privacy is under constant threat as our personal data and traceable online/offline activities are frequently collected, processed and transferred by many software applications. Privacy attacks are often formed by exploiting vulnerabilities found in those software applications. The Common Weakness Enumeration (CWE) and Common Vulnerabilities and Exposures (CVE) systems are currently the main sources that software engineers rely on for understanding and preventing publicly disclosed software vulnerabilities. However, our study on all 922 weaknesses in the CWE and 156,537 vulnerabilities registered in the CVE to date has found a very small coverage of privacy-related vulnerabilities in both systems, only 4.45% in CWE and 0.1% in CVE. These also cover only a small number of areas of privacy threats that have been raised in existing privacy software engineering research, privacy regulations and frameworks, and relevant reputable organisations. The actionable insights generated from our study led to the introduction of 11 new common privacy weaknesses to supplement the CWE system, making it become a source for both security and privacy vulnerabilities. Pattaraporn Sangaroonsilp, Khanh Hoa Dam, Aditya Ghose |
ICSE | 2 |
| 2023 | A normative approach for resilient multiagent systems
Geeta Mahala, Özgür Kafali, Khanh Hoa Dam, Aditya Ghose, Munindar P. Singh |
Auton. Agents Multi Agent Syst. | 3 |
| 2023 | An empirical study of automated privacy requirements classification in issue reportsabstractAbstract The recent advent of data protection laws and regulations has emerged to protect privacy and personal information of individuals. As the cases of privacy breaches and vulnerabilities are rapidly increasing, people are aware and more concerned about their privacy. These bring a significant attention to software development teams to address privacy concerns in developing software applications. As today’s software development adopts an agile, issue-driven approach, issues in an issue tracking system become a centralised pool that gathers new requirements, requests for modification and all the tasks of the software project. Hence, establishing an alignment between those issues and privacy requirements is an important step in developing privacy-aware software systems. This alignment also facilitates privacy compliance checking which may be required as an underlying part of regulations for organisations. However, manually establishing those alignments is labour intensive and time consuming. In this paper, we explore a wide range of machine learning and natural language processing techniques which can automatically classify privacy requirements in issue reports. We employ six popular techniques namely Bag-of-Words (BoW), N-gram Inverse Document Frequency (N-gram IDF), Term Frequency-Inverse Document Frequency (TF-IDF), Word2Vec, Convolutional Neural Network (CNN) and Bidirectional Encoder Representations from Transformers (BERT) to perform the classification on privacy-related issue reports in Google Chrome and Moodle projects. The evaluation showed that BoW, N-gram IDF, TF-IDF and Word2Vec techniques are suitable for classifying privacy requirements in those issue reports. In addition, N-gram IDF is the best performer in both projects. Pattaraporn Sangaroonsilp, Morakot Choetkiertikul, Khanh Hoa Dam, Aditya Ghose |
Autom. Softw. Eng. | 3 |
| 2023 | A taxonomy for mining and classifying privacy requirements in issue reports
Pattaraporn Sangaroonsilp, Khanh Hoa Dam, Morakot Choetkiertikul, Chaiyong Ragkhitwetsagul, Aditya Ghose |
Inf. Softw. Technol. | 2 |
| 2023 | Egalitarian Transient Service Composition in Crowdsourced IoT EnvironmentabstractThe Crowdsourced IoT Service (CIS) market is inherently different from other service markets, e.g., web services and cloud. The CIS market is dominated by transient services as both consumers and providers are dynamic in space and time. Consumer requests are usually long-term and demand continuity in service provision. We propose a novel egalitarian transient service composition framework from the CIS market perspective. We apply a Dynamic Bayesian Network to model the dynamic service provision behavior of the providers. The proposed framework transforms the composition of transient services into a multi-objective temporal optimization, i.e., providing continuous services to the maximum number of consumers, and minimizing the consumers’ cost of service usages over a long-term period. We incorporate a Pareto-based genetic algorithm to enable the fair distribution of services among the consumers. Experimental results prove the efficiency of the proposed approach in terms of continuous availability of service as well as fair distribution among consumers. Swasti Khurana, Novarun Deb, Sajib Mistry, Aditya Ghose, Aneesh Krishna, Khanh Hoa Dam |
IEEE Trans. Serv. Comput. | 6 |
| 2022 | Goal-Oriented Coordination with Cumulative Goals
Aditya Ghose, Khanh Hoa Dam, Shunichiro Tomura, Dean Philp, Angela Consoli |
PRIMA | 2 |
| 2022 | DITURIA: A Framework for Decision Coordination Among Multiple Agents
Helena Ibro, Geeta Mahala, Simon Pulawski, Steven Harvey, Alexis A. Miller, Aditya Ghose, Khanh Hoa Dam |
PRIMA | 7 |
| 2022 | A framework for conditional statement technical debt identification and descriptionabstractAbstract Technical Debt occurs when development teams favour short-term operability over long-term stability. Since this places software maintainability at risk, technical debt requires early attention to avoid paying for accumulated interest. Most of the existing work focuses on detecting technical debt using code comments, known as Self-Admitted Technical Debt (SATD). However, there are many cases where technical debt instances are not explicitly acknowledged but deeply hidden in the code. In this paper, we propose a framework that caters for the absence of SATD comments in code. Our Self-Admitted Technical Debt Identification and Description (SATDID) framework determines if technical debt should be self-admitted for an input code fragment. If that is the case, SATDID will automatically generate the appropriate descriptive SATD comment that can be attached with the code. While our approach is applicable in principle to any type of code fragments, we focus in this study on technical debt hidden in conditional statements, one of the most TD-carrying parts of code. We explore and evaluate different implementations of SATDID. The evaluation results demonstrate the applicability and effectiveness of our framework over multiple benchmarks. Comparing with the results from the benchmarks, our approach provides at least 21.35, 59.36, 31.78, and 583.33% improvements in terms of Precision, Recall, F-1, and Bleu-4 scores, respectively. In addition, we conduct a human evaluation to the SATD comments generated by SATDID. In 1-5 and 0–5 scales for Acceptability and Understandability, the total means achieved by our approach are 3.128 and 3.172, respectively. Abdulaziz Alhefdhi, Khanh Hoa Dam, Yusuf Sulistyo Nugroho, Hideaki Hata, Takashi Ishio, Aditya Ghose |
Autom. Softw. Eng. | 2 |
| 2022 | An Empirical Study of Model-Agnostic Techniques for Defect Prediction ModelsabstractSoftware analytics have empowered software organisations to support a wide range of improved decision-making and policy-making. However, such predictions made by software analytics to date have not been explained and justified. Specifically, current defect prediction models still fail to explain why models make such a prediction and fail to uphold the privacy laws in terms of the requirement to explain any decision made by an algorithm. In this paper, we empirically evaluate three model-agnostic techniques, i.e., two state-of-the-art Local Interpretability Model-agnostic Explanations technique (LIME) and BreakDown techniques, and our improvement of LIME with Hyper Parameter Optimisation (LIME-HPO). Through a case study of 32 highly-curated defect datasets that span across 9 open-source software systems, we conclude that (1) model-agnostic techniques are needed to explain individual predictions of defect models; (2) instance explanations generated by model-agnostic techniques are mostly overlapping (but not exactly the same) with the global explanation of defect models and reliable when they are re-generated; (3) model-agnostic techniques take less than a minute to generate instance explanations; and (4) more than half of the practitioners perceive that the contrastive explanations are necessary and useful to understand the predictions of defect models. Since the implementation of the studied model-agnostic techniques is available in both Python and R, we recommend model-agnostic techniques be used in the future. Jirayus Jiarpakdee, Chakkrit Tantithamthavorn, Khanh Hoa Dam, John C. Grundy |
IEEE Trans. Software Eng. | 3 |
| 2021 | A Fuzzy-Based Requirement Selection Method for Considering Value Dependencies in Software Release PlanningabstractRequirement selection is an essential component of software release planning, which finds, for a given budget, an optimal subset of the requirements with the highest value. However, due to the dependencies among software requirements, selecting or ignoring a requirement may impact the values of others. But such Value Dependencies are imprecise and hard to capture; they have been ignored by the existing requirement selection methods, increasing the risk of value loss in software projects. To address this, we have proposed a fuzzy-based optimization method with two main components: (i) a fuzzy-based technique for modeling value dependencies and capturing their imprecision, and (ii) an Integer Linear Programming (ILP) model that takes into account value dependencies in software requirement selection. The scalability and effectiveness of the method in mitigating value loss are demonstrated through simulations. Davoud Mougouei, Aditya Ghose, Khanh Hoa Dam, Mahdi Fahmideh, David M. W. Powers |
FUZZ-IEEE | 3 |
| 2021 | Cross-Silo Process Mining with Federated Learning
Muhammad Asjad Khan, Aditya Ghose, Khanh Hoa Dam |
ICSOC | 3 |
| 2021 | DeepProcess: Supporting Business Process Execution Using a MANN-Based Recommender System
Muhammad Asjad Khan, Hung Le 0002, Kien Do, Truyen Tran 0001, Aditya Ghose, Khanh Hoa Dam, Renuka Sindhgatta |
ICSOC | 6 |
| 2021 | Automatically recommending components for issue reports using deep learning
Morakot Choetkiertikul, Khanh Hoa Dam, Truyen Tran 0001, Trang Pham, Chaiyong Ragkhitwetsagul, Aditya Ghose |
Empir. Softw. Eng. | 2 |
| 2021 | Automatic Feature Learning for Predicting Vulnerable Software ComponentsabstractCode flaws or vulnerabilities are prevalent in software systems and can potentially cause a variety of problems including deadlock, hacking, information loss and system failure. A variety of approaches have been developed to try and detect the most likely locations of such code vulnerabilities in large code bases. Most of them rely on manually designing code features (e.g., complexity metrics or frequencies of code tokens) that represent the characteristics of the potentially problematic code to locate. However, all suffer from challenges in sufficiently capturing both semantic and syntactic representation of source code, an important capability for building accurate prediction models. In this paper, we describe a new approach, built upon the powerful deep learning Long Short Term Memory model, to automatically learn both semantic and syntactic features of code. Our evaluation on 18 Android applications and the Firefox application demonstrates that the prediction power obtained from our learned features is better than what is achieved by state of the art vulnerability prediction models, for both within-project prediction and cross-project prediction. Khanh Hoa Dam, Truyen Tran 0001, Trang Pham, Shien Wee Ng, John C. Grundy, Aditya Ghose |
IEEE Trans. Software Eng. | 1 |
| 2020 | Designing Optimal Robotic Process Automation Architectures
Geeta Mahala, Renuka Sindhgatta, Khanh Hoa Dam, Aditya Ghose |
ICSOC | 3 |
| 2020 | G2I: A principled approach to leveraging the goal-to-information nexus in BDI agentsabstractIn a number of settings, information requirements must be guided by the current goals of an agent. This paper shows how information requirements can be systematically derived from goals by examining the executability conditions of plans and actions that help achieve that goal. We extend the standard BDI framework by annotating beliefs, plans and actions with probabilities and propose the use of information-seeking actions to meet the information requirements. We show how agents can use such information-seeking actions, which can be expensive to execute (often at the cost of other mission-critical actions), to increase the expected utility of their goals, and thereby present a novel variant of the standard BDI framework. Kinzang Chhogyal, Angela Consoli, Aditya Ghose, Khanh Hoa Dam |
KES | 4 |
| 2020 | Abductive Design of BDI Agent-Based Digital Twins of Organizations
Ahmad Alelaimat, Aditya Ghose, Khanh Hoa Dam |
PRIMA | 3 |
| 2019 | GameOfFlows: Process Instance Adaptation in Complex, Dynamic and Potentially Adversarial Domains
Yingzhi Gou, Aditya Ghose, Khanh Hoa Dam |
CAiSE | 3 |
| 2019 | A Value-based Trust Assessment Model for Multi-agent SystemsabstractAn agent's assessment of its trust in another agent is commonly taken to be a measure of the reliability/predictability of the latter's actions. It is based on the trustor's past observations of the behaviour of the trustee and requires no knowledge of the inner-workings of the trustee. However, in situations that are new or unfamiliar, past observations are of little help in assessing trust. In such cases, knowledge about the trustee can help. A particular type of knowledge is that of values - things that are important to the trustor and the trustee. In this paper, based on the premise that the more values two agents share, the more they should trust one another, we propose a simple approach to trust assessment between agents based on values, taking into account if agents trust cautiously or boldly, and if they depend on others in carrying out a task. Kinzang Chhogyal, Abhaya C. Nayak, Aditya Ghose, Khanh Hoa Dam |
IJCAI | 4 |
| 2019 | Lessons learned from using a deep tree-based model for software defect prediction in practiceabstractDefects are common in software systems and cause many problems for software users. Different methods have been developed to make early prediction about the most likely defective modules in large codebases. Most focus on designing features (e.g. complexity metrics) that correlate with potentially defective code. Those approaches however do not sufficiently capture the syntax and multiple levels of semantics of source code, a potentially important capability for building accurate prediction models. In this paper, we report on our experience of deploying a new deep learning tree-based defect prediction model in practice. This model is built upon the tree-structured Long Short Term Memory network which directly matches with the Abstract Syntax Tree representation of source code. We discuss a number of lessons learned from developing the model and evaluating it on two datasets, one from open source projects contributed by our industry partner Samsung and the other from the public PROMISE repository. Khanh Hoa Dam, Trang Pham, Shien Wee Ng, Truyen Tran 0001, John C. Grundy, Aditya Ghose, Taeksu Kim, Chul-Joo Kim |
MSR | 1 |
| 2019 | DeepJIT: an end-to-end deep learning framework for just-in-time defect predictionabstractSoftware quality assurance efforts often focus on identifying defective code. To find likely defective code early, change-level defect prediction - aka. Just-In-Time (JIT) defect prediction - has been proposed. JIT defect prediction models identify likely defective changes and they are trained using machine learning techniques with the assumption that historical changes are similar to future ones. Most existing JIT defect prediction approaches make use of manually engineered features. Unlike those approaches, in this paper, we propose an end-to-end deep learning framework, named DeepJIT, that automatically extracts features from commit messages and code changes and use them to identify defects. Experiments on two popular software projects (i.e., QT and OPENSTACK) on three evaluation settings (i.e., cross-validation, short-period, and long-period) show that the best variant of DeepJIT (DeepJIT-Combined), compared with the best performing state-of-the-art approach, achieves improvements of 10.36-11.02% for the project QT and 9.51-13.69% for the project OPENSTACK in terms of the Area Under the Curve (AUC). Thong Hoang, Khanh Hoa Dam, Yasutaka Kamei, David Lo 0001, Naoyasu Ubayashi |
MSR | 2 |
| 2019 | A Deep Learning Model for Estimating Story PointsabstractAlthough there has been substantial research in software analytics for effort estimation in traditional software projects, little work has been done for estimation in agile projects, especially estimating the effort required for completing user stories or issues. Story points are the most common unit of measure used for estimating the effort involved in completing a user story or resolving an issue. In this paper, we propose a prediction model for estimating story points based on a novel combination of two powerful deep learning architectures: long short-term memory and recurrent highway network. Our prediction system is end-to-end trainable from raw input data to prediction outcomes without any manual feature engineering. We offer a comprehensive dataset for story points-based estimation that contains 23,313 issues from 16 open source projects. An empirical evaluation demonstrates that our approach consistently outperforms three common baselines (Random Guessing, Mean, and Median methods) and six alternatives (e.g., using Doc2Vec and Random Forests) in Mean Absolute Error, Median Absolute Error, and the Standardized Accuracy. Morakot Choetkiertikul, Khanh Hoa Dam, Truyen Tran 0001, Trang Pham, Aditya Ghose, Tim Menzies |
IEEE Trans. Software Eng. | 2 |
| 2018 | Multi-Objective Iteration Planning in Agile DevelopmentabstractAn agile software project typically has a number of iterations (e.g. sprints in Scrum), in each of which the development team designs, implements, tests and delivers a distinct product increment. An important activity in agile development is iteration planning where the team needs to decide what should be done (in terms of issues or user stories) for the upcoming iteration. In this paper, we propose a multi-objective search-based approach to support the team in making such a decision. Our approach employs evolutionary techniques to iteratively generate candidate selections of issues for a given iteration, and search for the optimal selection(s). The search is guided simultaneously by two objectives: maximizing the business value which the team delivers in the iteration while maximizing the alignment with regard to the iteration's original goal. Our evaluation of 233 iterations from six large open source projects demonstrates the effectiveness of our approach. Wisam Haitham Abbood Al-Zubaidi, Khanh Hoa Dam, Morakot Choetkiertikul, Aditya Ghose |
APSEC | 2 |
| 2018 | Leveraging Regression Algorithms for Process Performance Predictions
Karthikeyan Ponnalagu, Aditya Ghose, Khanh Hoa Dam |
ICSOC | 3 |
| 2018 | Predicting Delivery Capability in Iterative Software DevelopmentabstractIterative software development has become widely practiced in industry. Since modern software projects require fast, incremental delivery for every iteration of software development, it is essential to monitor the execution of an iteration, and foresee a capability to deliver quality products as the iteration progresses. This paper presents a novel, data-driven approach to providing automated support for project managers and other decision makers in predicting delivery capability for an ongoing iteration. Our approach leverages a history of project iterations and associated issues, and in particular, we extract characteristics of previous iterations and their issues in the form of features. In addition, our approach characterizes an iteration using a novel combination of techniques including feature aggregation statistics, automatic feature learning using the Bag-of-Words approach, and graph-based complexity measures. An extensive evaluation of the technique on five large open source projects demonstrates that our predictive models outperform three common baseline methods in Normalized Mean Absolute Error and are highly accurate in predicting the outcome of an ongoing iteration. Morakot Choetkiertikul, Khanh Hoa Dam, Truyen Tran 0001, Aditya Ghose, John C. Grundy |
IEEE Trans. Software Eng. | 2 |
| 2017 | Leveraging Game-Tree Search for Robust Process Enactment
Yingzhi Gou, Aditya Ghose, Khanh Hoa Dam |
CAiSE | 3 |
| 2017 | Mining Goal Refinement Patterns: Distilling Know-How from Data
Metta Santiputri, Novarun Deb, Muhammad Asjad Khan, Aditya Ghose, Khanh Hoa Dam, Nabendu Chaki |
ER | 5 |
| 2017 | Goal Orchestrations: Modelling and Mining Flexible Business Processes
Metta Santiputri, Aditya Ghose, Khanh Hoa Dam, Suman Roy 0001 |
ER | 3 |
| 2017 | Mining task post-conditions: Automating the acquisition of process semantics
Metta Santiputri, Aditya Ghose, Khanh Hoa Dam |
Data Knowl. Eng. | 3 |
| 2017 | Predicting the delay of issues with due dates in software projects
Morakot Choetkiertikul, Khanh Hoa Dam, Truyen Tran 0001, Aditya Ghose |
Empir. Softw. Eng. | 2 |
| 2016 | Context-Aware Analysis of Past Process Executions to Aid Resource Allocation Decisions
Renuka Sindhgatta, Aditya Ghose, Khanh Hoa Dam |
CAiSE | 3 |
| 2016 | Context-Aware Recommendation of Task Allocations in Service Systems
Renuka Sindhgatta, Aditya Ghose, Khanh Hoa Dam |
ICSOC | 3 |
| 2016 | Externalization of software behavior by the mining of normsabstractOpen Source Software Development (OSSD) often suffers from conflicting views and actions due to the perceived flat and open ecology of an open source community. This often manifests itself as a lack of codified knowledge that is easily accessible for community members. How decisions are made and expectations of a software system are often described in detail through the many forms of social communications that take place within a community. These social interactions form norms which are influential in dictating what behaviors are expected in a community and of the system. In this paper, we provide a tool which mines these social interactions (in the form of bug reports) and extract norms of the system, externalizing this information into a codified form that allows others within the community to be aware of without having witnessed the social interactions. Daniel Avery, Khanh Hoa Dam, Bastin Tony Roy Savarimuthu, Aditya Ghose |
MSR | 2 |
| 2016 | Analyzing Topics and Trends in the PRIMA Literature
Khanh Hoa Dam, Aditya Ghose |
PRIMA | 1 |
| 2016 | DeepSoft: a vision for a deep model of softwareabstractAlthough software analytics has experienced rapid growth as a research area, it has not yet reached its full potential for wide industrial adoption. Most of the existing work in software analytics still relies heavily on costly manual feature engineering processes, and they mainly address the traditional classification problems, as opposed to predicting future events. We present a vision for DeepSoft, an end-to-end generic framework for modeling software and its development process to predict future risks and recommend interventions. DeepSoft, partly inspired by human memory, is built upon the powerful deep learning-based Long Short Term Memory architecture that is capable of learning long-term temporal dependencies that occur in software evolution. Such deep learned patterns of software can be used to address a range of challenging problems such as code and task recommendation and prediction. DeepSoft provides a new approach for research into modeling of source code, risk prediction and mitigation, developer modeling, and automatically generating code patches from bug reports. Khanh Hoa Dam, Truyen Tran 0001, John C. Grundy, Aditya Ghose |
SIGSOFT FSE | 1 |
| 2016 | Consistent merging of model versions
Khanh Hoa Dam, Alexander Egyed, Michael Winikoff, Alexander Reder, Roberto Erick Lopez-Herrejon |
J. Syst. Softw. | 1 |
| 2015 | Goal-Aligned Categorization of Instance Variants in Knowledge-Intensive Processes
Karthikeyan Ponnalagu, Aditya Ghose, Nanjangud C. Narendra, Khanh Hoa Dam |
BPM | 4 |
| 2015 | Mining Process Task Post-Conditions
Metta Santiputri, Aditya Ghose, Khanh Hoa Dam, Xiong Wen |
ER | 3 |
| 2015 | Learning Relationships Between the Business Layer and the Application Layer in ArchiMate Models
Ayu Saraswati, Chee Fon Chang, Aditya Ghose, Khanh Hoa Dam |
ER | 4 |
| 2015 | Mining Software Repositories for Social NormsabstractSocial norms facilitate coordination and cooperation among individuals, thus enable smoother functioning of social groups such as the highly distributed and diverse open source software development (OSSD) communities. In these communities, norms are mostly implicit and hidden in huge records of human-interaction information such as emails, discussions threads, bug reports, commit messages and even source code. This paper aims to introduce a new line of research on extracting social norms from the rich data available in software repositories. Initial results include a study of coding convention violations in JEdit, Argo UML and Glassfish projects. It also presents a new life-cycle model for norms in OSSD communities and demonstrates how a number of norms extracted from the Python development community follow this life-cycle model. Khanh Hoa Dam, Bastin Tony Roy Savarimuthu, Daniel Avery, Aditya Ghose |
ICSE (2) | 1 |
| 2015 | Predicting Delays in Software Projects Using Networked Classification (T)abstractSoftware projects have a high risk of cost and schedule overruns, which has been a source of concern for the software engineering community for a long time. One of the challenges in software project management is to make reliable prediction of delays in the context of constant and rapid changes inherent in software projects. This paper presents a novel approach to providing automated support for project managers and other decision makers in predicting whether a subset of software tasks (among the hundreds to thousands of ongoing tasks) in a software project have a risk of being delayed. Our approach makes use of not only features specific to individual software tasks (i.e. local data) -- as done in previous work -- but also their relationships (i.e. networked data). In addition, using collective classification, our approach can simultaneously predict the degree of delay for a group of related tasks. Our evaluation results show a significant improvement over traditional approaches which perform classification on each task independently: achieving 46% -- 97% precision (49% improved), 46% -- 97% recall (28% improved), 56% -- 75% F-measure (39% improved), and 78% -- 95% Area Under the ROC Curve (16% improved). Morakot Choetkiertikul, Khanh Hoa Dam, Truyen Tran 0001, Aditya Ghose |
ASE | 2 |
| 2015 | Characterization and Prediction of Issue-Related Risks in Software ProjectsabstractIdentifying risks relevant to a software project and planning measures to deal with them are critical to the success of the project. Current practices in risk assessment mostly rely on high-level, generic guidance or the subjective judgements of experts. In this paper, we propose a novel approach to risk assessment using historical data associated with a software project. Specifically, our approach identifies patterns of past events that caused project delays, and uses this knowledge to identify risks in the current state of the project. A set of risk factors characterizing “risky” software tasks (in the form of issues) were extracted from five open source projects: Apache, Duraspace, JBoss, Moodle, and Spring. In addition, we performed feature selection using a sparse logistic regression model to select risk factors with good discriminative power. Based on these risk factors, we built predictive models to predict if an issue will cause a project delay. Our predictive models are able to predict both the risk impact (i.e. the extend of the delay) and the likelihood of a risk occurring. The evaluation results demonstrate the effectiveness of our predictive models, achieving on average 48%-81% precision, 23%-90% recall, 29%-71% F-measure, and 70%-92% Area Under the ROC Curve. Our predictive models also have low error rates: 0.39-0.75 for Macro-averaged Mean Cost-Error and 0.7-1.2 for Macro-averaged Mean Absolute Error. Morakot Choetkiertikul, Khanh Hoa Dam, Truyen Tran 0001, Aditya Ghose |
MSR | 2 |
| 2014 | A CMMI-Based Automated Risk Assessment FrameworkabstractRisk assessment is crucial to the increase of software development project success. Current risk assessment approaches provide only a rough guide. Risk assessment experts and domain experts are required in conducting risk assessments in software projects. Therefore, traditional risk assessment approaches require extra activities besides development tasks, and possibly leading to extra costs. We believe that an effective risk assessment approach should be transparently embedded in software development process. This paper aims to present an automated risk assessment framework using CMMI and risk taxnomy as a guidance to develop a risk assessment model. A pragmatic approach will be applied as a basis in building this suggested risk prediction model and the case studies of our practice. These studies are considered as our proof of concept. Morakot Choetkiertikul, Khanh Hoa Dam, Aditya Ghose, Thanwadee Sunetnanta |
APSEC (2) | 2 |
| 2014 | Inconsistency Resolution in Merging Versions of Architectural ModelsabstractState-of-the-art optimistic model versioning systems, which are critical to enable efficient team-based development of architectural models, are able to detect and help resolve basic conflicts arising during the merging of model versions. However, it is often overlooked that model merging may also cause severe syntactical and semantic inconsistencies. In this paper, we propose an approach to guide the resolution of inconsistencies detected in a merged architectural model. Our approach automatically finds and presents to the software architects all solutions for resolving all inconsistencies arisen during the merging of model versions. For inconsistencies that pre-exist in the model, our approach is able to suggest exactly which model elements should be changed to resolve them. Our approach is built upon a repair generation which can quickly derive resolutions for an inconsistency by examining its static and dynamic structure and forming concrete repair actions from changes in the versions to be merged. An empirical validation on a range of industrial models has demonstrated that our approach is scalable to both large models and large differences between model versions. Khanh Hoa Dam, Alexander Reder, Alexander Egyed |
WICSA | 1 |
| 2013 | Improving the Reactivity of BDI Agent Programs
Khanh Hoa Dam, Tiancheng Zhang 0006, Aditya Ghose |
PRIMA | 1 |
| 2013 | Towards Semantic Merging of Versions of BDI Agent Systems
Yingzhi Gou, Khanh Hoa Dam, Aditya Ghose |
PRIMA | 2 |
| 2013 | Supporting change impact analysis for intelligent agent systems
Khanh Hoa Dam, Aditya Ghose |
Sci. Comput. Program. | 1 |
| 2013 | Towards a next-generation AOSE methodology
Khanh Hoa Dam, Michael Winikoff |
Sci. Comput. Program. | 1 |
| 2012 | Maintaining Motivation Models (in BMM) in the Context of a (WSDL-S) Service Landscape
Konstantin Hoesch-Klohe, Aditya Ghose, Khanh Hoa Dam |
ICSOC | 3 |
| 2012 | Relationship-Preserving Change Propagation in Process Ecosystems
Tri A. Kurniawan, Aditya Ghose, Khanh Hoa Dam, Lam-Son Lê |
ICSOC | 3 |
| 2011 | Automated change impact analysis for agent systemsabstractIntelligent agent technology has evolved rapidly over the past few years along with the growing number of agent systems in various domains. Although a substantial amount of work in agent-oriented software engineering has provided methodologies for analysing, designing and implementing agent-based systems, recent studies have highlighted that there has been very little work on maintenance and evolution of agent-based systems. A critical issue in software maintenance and evolution is change impact analysis: determining the potential consequences of a proposed change. There has been a proliferation of techniques proposed to support change impact analysis of procedural or object-oriented systems, but to the best of our knowledge, no such an effort has been made for agent-based software. In this paper, we fill this gap by proposing a framework to support change impact analysis for agent systems. At the core of our framework is the taxonomy of atomic changes which can precisely capture semantic differences between versions of an agent system. We also present a change impact model in the form of an intra-agent dependency graph that represents various dependencies within an agent system. An algorithm to compute the set of entities impacted by a change is also presented. The proposed techniques have been implemented in AgentCIA, a change impact analysis plugin for Jason, one of the most well-known agent programming platforms. Khanh Hoa Dam, Aditya Ghose |
ICSM | 1 |
| 2011 | An agent-oriented approach to change propagation in software maintenance
Khanh Hoa Dam, Michael Winikoff |
Auton. Agents Multi Agent Syst. | 1 |
| 2010 | Supporting Change Propagation in the Maintenance and Evolution of Service-Oriented ArchitecturesabstractAs Service-Oriented Architecture (SOA) continues to be broadly adopted, the maintenance and evolution of service-oriented systems become a growing issue. Maintenance and evolution are inevitable activities since almost all systems that are useful and successful stimulate user-generated requests for change and improvement. A critical issue in the evolution of SOA is change propagation: given a set of primary changes that have been made to the SOA model, what additional secondary changes are needed to maintain consistency across multiple levels of the SOA models. This paper presents how an existing framework can be applied to effectively support change propagation within a SOA model. We also propose to extend this framework with a minimal modification strategy that helps select change options in a manner that accommodates the structural and semantic dimensions of SOA models. Khanh Hoa Dam, Aditya Ghose |
APSEC | 1 |
| 2010 | Supporting Change Propagation in the Evolution of Enterprise ArchitecturesabstractEnterprise Architecture (EA) models the whole vision of an organisation in various aspects regarding both business processes and information technology resources. As the organisation grows, the architecture governing its systems and processes must also evolve to meet with the demands of the business environment. In this context, a critical issue is change propagation: given a set of primary changes that have been made to the EA model, what additional secondary changes are needed to maintain consistency across multiple levels of the EA. This paper proposes an enterprise architectural description language, namely Change Aware Hierarchical EA, integrated with a framework to support change propagation within an EA model. The core part of our change propagation framework is a new method for generating interactive repair plans from Alloy consistency rules that constrain the EA model. Khanh Hoa Dam, Lam-Son Lê, Aditya Ghose |
EDOC | 1 |
| 2010 | Supporting change propagation in UML modelsabstractA critical issue in software maintenance and evolution is change propagation: given a primary change that is made in order to meet a new or changed requirement, what additional, secondary, changes are needed? We have previously developed techniques for effectively supporting change propagation within design models of intelligent agent systems. In this paper, we propose how this approach is applied to support change propagation within UML design models. Our approach offers a number of advantages in terms of saving substantial time writing hard-coded rules, ensuring soundness and completeness, and at the same time capturing the cascading nature of change propagation. We will also present and discuss the results of an evaluation performed to assess the scalability of our approach. Khanh Hoa Dam, Michael Winikoff |
ICSM | 1 |
| 2010 | Agent-Based Development for Business Processes
Khanh Hoa Dam, Aditya Ghose |
PRIMA | 1 |
| 2010 | An Agent-Oriented Approach to Service Analysis and Design
Khanh Hoa Dam, Aditya Ghose |
PRIMA | 1 |
| 2010 | Business Rules Discovery from Process Design RepositoriesabstractTraditional process mining approaches focus on extracting process constraints or business rules from repositories of process instances. In this context, process designs or process models tend to be overlooked although they contain information that are valuable for the process of discovering business rules. This paper will propose an alternative approach to process mining in terms of using process designs as the mining resources. We propose a number of techniques for extracting business rules from repositories of business process designs or models, leveraging the well-known Apriori algorithm. Such business rules are then used as a prior knowledge for further analysing, verifying, and modifying process designs. Jantima Polpinij, Aditya Ghose, Khanh Hoa Dam |
SERVICES | 3 |
| 2003 | Comparing particle swarms for tracking extrema in dynamic environmentsabstractThis work presents a comparative study of particle swarm models on their abilities to track extrema in dynamic environments. A standard PSO, two randomized PSOs, and a fine-grained PSO are evaluated in non-trivial multimodal dynamic environments involving small constant step changes, different large step changes, and chaotic step changes of the extrema. DF1 proposed by Morrison and De Jong is used to generate these three types of dynamics (1999). Our results indicate that PSO and its variants are able to perform reasonably well in a 2-dimensional variable space, whereas perform well to a less extent in a 10-dimensional variable space. It is also found that the fine-grained PSO is able to outperform all other PSO variants in the 10-dimensional variable space, likely due to its ability in maintaining better population diversity. Xiaodong Li 0001, Khanh Hoa Dam |
IEEE Congress on Evolutionary Computation | 2 |